Compare commits
251
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
a18060e9c0 | ||
|
|
3e61fde06d | ||
|
|
f0ac5b48ce | ||
|
|
1141ebcd49 | ||
|
|
82545548e7 | ||
|
|
aace66e8d3 | ||
|
|
b1627e9ba0 | ||
|
|
a4fa16ae69 | ||
|
|
80fbc779d9 | ||
|
|
07eb682358 | ||
|
|
3ad817eeff | ||
|
|
26dc1722f4 | ||
|
|
c4a4f1f25c | ||
|
|
8296153a38 | ||
|
|
bf18712a6c | ||
|
|
c0643af118 | ||
|
|
80c4ea8966 | ||
|
|
e75448a118 | ||
|
|
720ab81bf8 | ||
|
|
9baa0a10af | ||
|
|
151c9d2e30 | ||
|
|
05e8c64e47 | ||
|
|
7f772bccbb | ||
|
|
c49336ba21 | ||
|
|
5cadbd40a9 | ||
|
|
a27d76ae1d | ||
|
|
e836a73a8a | ||
|
|
e521cf5976 | ||
|
|
13b65e19ec | ||
|
|
9bda3a52fa | ||
|
|
f69cdec9d4 | ||
|
|
22becd6859 | ||
|
|
47890ca913 | ||
|
|
fa8db4c7e8 | ||
|
|
d17524cd67 | ||
|
|
d50492b86f | ||
|
|
0c29ab3a96 | ||
|
|
215bfd78d7 | ||
|
|
857b6da886 | ||
|
|
5521e498a9 | ||
|
|
346a93921e | ||
|
|
c1796bea17 | ||
|
|
bd909d5858 | ||
|
|
857c73a66a | ||
|
|
c4e4953f3d | ||
|
|
15f729ffea | ||
|
|
18e2397272 | ||
|
|
248fa5d100 | ||
|
|
c7db7d4126 | ||
|
|
abfe60d62d | ||
|
|
7819e5de50 | ||
|
|
3424852a39 | ||
|
|
8f684cedc0 | ||
|
|
955433ba81 | ||
|
|
4648175773 | ||
|
|
3247b9e61a | ||
|
|
b9c97ffac9 | ||
|
|
b9ebc872ad | ||
|
|
c57daaf861 | ||
|
|
43699b46db | ||
|
|
854320ac3f | ||
|
|
ad4d256d8f | ||
|
|
3ed8bd7be9 | ||
|
|
96aa8176cd | ||
|
|
5cbfa893a5 | ||
|
|
2961beb8fe | ||
|
|
b14bacbfc6 | ||
|
|
6e47730501 | ||
|
|
ac403b9cd3 | ||
|
|
7791ed74f2 | ||
|
|
6abce99b49 | ||
|
|
b76d0acec7 | ||
|
|
580032056e | ||
|
|
2755e41ff3 | ||
|
|
d5e623c7ba | ||
|
|
aa82321907 | ||
|
|
62e8c87d0e | ||
|
|
62c5a8a408 | ||
|
|
5f9380dfca | ||
|
|
aa5abfa911 | ||
|
|
c2fe6a6bf4 | ||
|
|
f6048f268b | ||
|
|
2d04c448c2 | ||
|
|
c59b38775b | ||
|
|
f5a76cf88a | ||
|
|
1184d6b787 | ||
|
|
811f2b8898 | ||
|
|
4c842f5d70 | ||
|
|
b4ef42d947 | ||
|
|
ecd42ed484 | ||
|
|
c9af5481e7 | ||
|
|
a2ff2a1102 | ||
|
|
4d4cdd6ea7 | ||
|
|
70c988e702 | ||
|
|
0790f8dfd3 | ||
|
|
9c3a1c5be1 | ||
|
|
cbb0f11288 | ||
|
|
9dad61f508 | ||
|
|
971ae01caf | ||
|
|
397a400d57 | ||
|
|
2c6739ad76 | ||
|
|
cec9a98305 | ||
|
|
de7fb2c936 | ||
|
|
01988305a8 | ||
|
|
328e570309 | ||
|
|
cf5a790ea8 | ||
|
|
4d3c85fd06 | ||
|
|
28fe7c43a9 | ||
|
|
72553cb414 | ||
|
|
20a95da487 | ||
|
|
c7e585e21d | ||
|
|
5fa8b7412e | ||
|
|
b8e554dac7 | ||
|
|
765a8923a4 | ||
|
|
fa0e8d7d97 | ||
|
|
13d64e0000 | ||
|
|
a9b275abbb | ||
|
|
92c06ac8dc | ||
|
|
168a37542b | ||
|
|
edd9d63f5e | ||
|
|
43df08b52a | ||
|
|
94f71eea19 | ||
|
|
17ede3c6aa | ||
|
|
ae6e9256c6 | ||
|
|
abb5910d2f | ||
|
|
5450ec786f | ||
|
|
24a6ab3d1e | ||
|
|
ac3a557566 | ||
|
|
508a1c02da | ||
|
|
55d515d41f | ||
|
|
f6dbfd3625 | ||
|
|
f378953982 | ||
|
|
daf760220b | ||
|
|
a31eca65c3 | ||
|
|
2f90851c03 | ||
|
|
0a36b3fda9 | ||
|
|
72c4aa3895 | ||
|
|
4933c075b0 | ||
|
|
11ac4f50e6 | ||
|
|
35d93d7612 | ||
|
|
97a64c8a33 | ||
|
|
2010961d32 | ||
|
|
f21aef3cfa | ||
|
|
4298cd5de1 | ||
|
|
2bd25be712 | ||
|
|
bb9798e32c | ||
|
|
3ffa3f5318 | ||
|
|
58535890c4 | ||
|
|
ba9d98f7ce | ||
|
|
6907961ce0 | ||
|
|
d0b1f9694e | ||
|
|
1918da29be | ||
|
|
fbb6b0c180 | ||
|
|
ffe5dc14a8 | ||
|
|
d4bb8d344b | ||
|
|
ed722d55f8 | ||
|
|
311b1a7ec4 | ||
|
|
36b954d347 | ||
|
|
30857df5b2 | ||
|
|
faa508e87a | ||
|
|
82b5a606f7 | ||
|
|
56f3abdb36 | ||
|
|
089d4f3a80 | ||
|
|
f650bf892a | ||
|
|
1c89a5eeeb | ||
|
|
c04a3f083e | ||
|
|
72f0b4a258 | ||
|
|
8e7c7bbf24 | ||
|
|
1d0ec61c9d | ||
|
|
bb68fefe04 | ||
|
|
d829267f1c | ||
|
|
e0d23780d8 | ||
|
|
d1ec40f738 | ||
|
|
b6ef27cd2d | ||
|
|
2a55a0d265 | ||
|
|
d2c656533e | ||
|
|
02fd2de502 | ||
|
|
f79e5ebb5e | ||
|
|
0c8e29b05a | ||
|
|
edefc34a5b | ||
|
|
87a9f4eb25 | ||
|
|
0a2d654e68 | ||
|
|
fe310743a2 | ||
|
|
2b87a5a13b | ||
|
|
fd33fd05e1 | ||
|
|
ff7c57cf9c | ||
|
|
a2df2f242b | ||
|
|
2a8f897e61 | ||
|
|
abce381faa | ||
|
|
a415246adc | ||
|
|
1efa8a4b08 | ||
|
|
c839454a1f | ||
|
|
e4f2cff532 | ||
|
|
0414913bc7 | ||
|
|
dcc3b7403e | ||
|
|
52549f7b3a | ||
|
|
9309ff5a7f | ||
|
|
e690b058db | ||
|
|
90ccbfede4 | ||
|
|
fd0794d04d | ||
|
|
a05edc934c | ||
|
|
a56c326518 | ||
|
|
7b5b28c587 | ||
|
|
0faec2b02a | ||
|
|
99c31c1d4e | ||
|
|
e9f74f3f0f | ||
|
|
fa7b54f5ab | ||
|
|
928a1fdfff | ||
|
|
b58c20311c | ||
|
|
cf65ffdae5 | ||
|
|
2458ee1722 | ||
|
|
6794e66c4d | ||
|
|
34f73ba19f | ||
|
|
0758b9c5d7 | ||
|
|
02c079c893 | ||
|
|
3b0fc7a3e0 | ||
|
|
404d1172a6 | ||
|
|
3918a4b11a | ||
|
|
515c4a6496 | ||
|
|
17b4396460 | ||
|
|
c2a5645c55 | ||
|
|
9ee8c48fff | ||
|
|
1ebd73a309 | ||
|
|
59ec23d4a8 | ||
|
|
e0ad78af98 | ||
|
|
30b4e1dfb2 | ||
|
|
942e9a5ff8 | ||
|
|
3f7274d29f | ||
|
|
a7fe525bfc | ||
|
|
8e9c8ca4e6 | ||
|
|
1d6c73007e | ||
|
|
fa3eda5228 | ||
|
|
1905cac950 | ||
|
|
72a2750461 | ||
|
|
2c74080b78 | ||
|
|
3338d6f0fe | ||
|
|
a4779186a4 | ||
|
|
ff81295aa9 | ||
|
|
a0f54df2a6 | ||
|
|
26f685be0e | ||
|
|
1dd62a9bdc | ||
|
|
2180e77cf5 | ||
|
|
4e98ae6e56 | ||
|
|
1976fca809 | ||
|
|
28d3638952 | ||
|
|
07bafebf0d | ||
|
|
8fe255e38f | ||
|
|
800a9042a1 | ||
|
|
d9246ddae6 | ||
|
|
3f2b28d0ec | ||
|
|
587f183191 |
No files matched your search
@@ -31,3 +31,14 @@ Dockerfile
|
||||
*.key
|
||||
felis
|
||||
felis.exe
|
||||
|
||||
# macOS materializes extended attributes as ._<name> sidecars (BSD tar uploads,
|
||||
# Finder copies, network volumes) and leaves .DS_Store behind. Neither is
|
||||
# source, and one is actively harmful: a ._*.sql beside the migrations is
|
||||
# //go:embed-ed into the binary and makes every `felis migrate` fail
|
||||
# ("non-numeric version") — observed live on a Mac-staged tree. Same exposure
|
||||
# for any tree the other //go:embed patterns walk (deploy/, plugins/).
|
||||
._*
|
||||
**/._*
|
||||
.DS_Store
|
||||
**/.DS_Store
|
||||
@@ -0,0 +1,23 @@
|
||||
# The workflows pin every action to a commit SHA, and every Dockerfile pins its base images
|
||||
# by digest. This keeps those pins moving: Dependabot reads the "# vX.Y.Z" comment next to
|
||||
# each action SHA, and the tag in front of each image digest, and opens a PR that bumps both
|
||||
# together.
|
||||
version: 2
|
||||
updates:
|
||||
- package-ecosystem: github-actions
|
||||
directory: /
|
||||
schedule:
|
||||
interval: weekly
|
||||
- package-ecosystem: docker
|
||||
directories:
|
||||
- /
|
||||
- /deploy/limbo
|
||||
- /deploy/lobby
|
||||
- /deploy/paper
|
||||
schedule:
|
||||
interval: weekly
|
||||
# A new major is a runtime change (Paper 26.x needs Java 25, Limbo's jar is Java 21
|
||||
# bytecode), so only digests and minors are proposed; majors move by hand.
|
||||
ignore:
|
||||
- dependency-name: "*"
|
||||
update-types: ["version-update:semver-major"]
|
||||
+119
-6
@@ -14,10 +14,14 @@
|
||||
# PR is what asks for the answer.
|
||||
name: ci
|
||||
|
||||
#
|
||||
# release.yml calls this workflow (workflow_call) before it builds anything, so a tag passes
|
||||
# exactly these gates and there is one list of them.
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
workflow_call:
|
||||
|
||||
permissions:
|
||||
contents: read
|
||||
@@ -30,19 +34,79 @@ jobs:
|
||||
go:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
- uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
|
||||
- name: gofmt
|
||||
run: |
|
||||
unformatted=$(gofmt -l .)
|
||||
if [ -n "$unformatted" ]; then
|
||||
echo "gofmt needed on:"; echo "$unformatted"; exit 1
|
||||
fi
|
||||
- run: go vet ./...
|
||||
- run: go test ./...
|
||||
# -race: felis-api and the operator are mostly goroutines (watchers, the
|
||||
# registry pruner, the backup scheduler, the rate limiters).
|
||||
- run: go test -race ./...
|
||||
# The version is pinned here and bumped by hand; Dependabot does not read `go run`.
|
||||
- name: staticcheck
|
||||
run: go run honnef.co/go/tools/cmd/[email protected] ./...
|
||||
|
||||
# Separate from the go job so a newly published advisory reads as what it is. govulncheck
|
||||
# exits non-zero only for vulnerable code this module can actually reach, standard
|
||||
# library included: setup-go installs the newest patch of go.mod's Go line, so a finding
|
||||
# there means the Dockerfile's golang digest (which ships the release) needs a bump too.
|
||||
vuln:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
- uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
|
||||
- run: go run golang.org/x/vuln/cmd/[email protected] ./...
|
||||
|
||||
# The business stores' SQL against a real PostgreSQL (internal/pgint): the unit suites run
|
||||
# on fakes, and PGRepo drifted from them three times while those stayed green. 13 is the
|
||||
# oldest server a supported distribution installs (EL9), 18 the newest (Arch).
|
||||
pgint:
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
fail-fast: false
|
||||
matrix:
|
||||
postgres: ['13', '18']
|
||||
services:
|
||||
postgres:
|
||||
image: postgres:${{ matrix.postgres }}
|
||||
env:
|
||||
POSTGRES_USER: felis
|
||||
POSTGRES_PASSWORD: pgint
|
||||
POSTGRES_DB: felis_pgint
|
||||
ports:
|
||||
- 5432:5432
|
||||
options: >-
|
||||
--health-cmd "pg_isready -U felis -d felis_pgint"
|
||||
--health-interval 2s
|
||||
--health-timeout 5s
|
||||
--health-retries 30
|
||||
steps:
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
- uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5.6.0
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
|
||||
- run: go test -race -tags pgint -count=1 ./internal/pgint/
|
||||
env:
|
||||
FELIS_TEST_PG_URL: postgres://felis:pgint@localhost:5432/felis_pgint?sslmode=disable
|
||||
|
||||
shell:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
# bootstrap.sh is the only thing that ever runs on a fresh host, and nothing here can
|
||||
# run it — it wants root, a package manager and k3s. Syntax plus the extracted-block
|
||||
@@ -60,12 +124,22 @@ jobs:
|
||||
esac
|
||||
done
|
||||
|
||||
# A pinned release rather than the runner image's copy, so a runner update cannot
|
||||
# change what fails. Warnings and errors fail the job; style notes (info) do not.
|
||||
- name: shellcheck
|
||||
run: |
|
||||
curl -fsSL -o shellcheck.tar.xz \
|
||||
https://github.com/koalaman/shellcheck/releases/download/v0.11.0/shellcheck-v0.11.0.linux.x86_64.tar.xz
|
||||
echo "8c3be12b05d5c177a04c29e3c78ce89ac86f1595681cab149b65b97c4e227198 shellcheck.tar.xz" | sha256sum -c
|
||||
tar -xJf shellcheck.tar.xz
|
||||
./shellcheck-v0.11.0/shellcheck -S warning $(git ls-files '*.sh')
|
||||
|
||||
- run: sh deploy/bootstrap_test.sh
|
||||
|
||||
panel:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
# The Dockerfile's `FROM node:<major>` is the only place the panel's Node version is
|
||||
# declared — there is no .nvmrc and no engines field. Reading it here rather than
|
||||
@@ -78,7 +152,7 @@ jobs:
|
||||
[ -n "$version" ] || { echo "Dockerfile has no 'FROM ... node:<major>' line"; exit 1; }
|
||||
echo "version=${version}" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- uses: actions/setup-node@v4
|
||||
- uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
|
||||
with:
|
||||
node-version: ${{ steps.node.outputs.version }}
|
||||
cache: npm
|
||||
@@ -92,3 +166,42 @@ jobs:
|
||||
|
||||
- run: npm run typecheck
|
||||
working-directory: panel
|
||||
|
||||
plugins:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
# The other jobs never touch the Java layer: the plugin jars were only ever
|
||||
# compiled by bootstrap on a live host, and the three test mains under
|
||||
# plugins/*/test were run by hand. JDK 21 plus the Gradle major the plugin
|
||||
# Dockerfiles pin (8.14) is that same toolchain, in CI.
|
||||
- uses: actions/setup-java@cf277c60eb25467037889841efdb72551f06f6c3 # v4.9.1
|
||||
with:
|
||||
distribution: temurin
|
||||
java-version: '21'
|
||||
|
||||
- uses: gradle/actions/setup-gradle@ed408507eac070d1f99cc633dbcf757c94c7933a # v4.4.3
|
||||
with:
|
||||
gradle-version: '8.14'
|
||||
|
||||
- run: bash plugins/test.sh
|
||||
|
||||
mods:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
# The three loader mods (Minecraft 1.20.1 / 1.20.4, Java-17 lines) compile
|
||||
# through their vendored Gradle wrappers, which fetch their own Gradle. Until
|
||||
# this job nothing ever built them: no install path touches them, and their
|
||||
# gradlew scripts were committed without the exec bit, so the README's
|
||||
# one-liners failed on a fresh clone.
|
||||
- uses: actions/setup-java@cf277c60eb25467037889841efdb72551f06f6c3 # v4.9.1
|
||||
with:
|
||||
distribution: temurin
|
||||
java-version: '17'
|
||||
|
||||
- uses: gradle/actions/setup-gradle@ed408507eac070d1f99cc633dbcf757c94c7933a # v4.4.3
|
||||
|
||||
- run: bash plugins/test-mods.sh
|
||||
+108
-18
@@ -17,6 +17,16 @@
|
||||
# quietly ships a release whose panel is that placeholder. The Dockerfile runs the npm
|
||||
# build first, and is the same recipe bootstrap uses, so there is one way to build felis
|
||||
# rather than two that can drift.
|
||||
#
|
||||
# SHA256SUMS is a contract with bootstrap too: download_release_binary refuses a binary whose
|
||||
# hash is not listed there, BEFORE it runs it. A release without the file installs by source
|
||||
# build instead.
|
||||
#
|
||||
# The write token never meets the test suite: `gates` (ci.yml) and `build` run the tests,
|
||||
# Gradle and the Docker build (each of which executes third-party code) with a read-only
|
||||
# token, and `build` hands the binaries over as a workflow artifact; `publish` holds contents:write and runs only
|
||||
# pinned actions and gh. Every action is pinned to a commit SHA (the tag in the trailing
|
||||
# comment is for humans); .github/dependabot.yml proposes the bumps.
|
||||
name: release
|
||||
|
||||
on:
|
||||
@@ -24,28 +34,29 @@ on:
|
||||
tags: ['v*']
|
||||
|
||||
permissions:
|
||||
contents: write # gh release create/upload
|
||||
contents: read
|
||||
|
||||
jobs:
|
||||
release:
|
||||
# A tag that ships red is worse than a tag that fails to ship. These are ci.yml's gates,
|
||||
# called rather than copied: Go (race, vet, staticcheck), govulncheck, the PostgreSQL
|
||||
# contract suite, shellcheck and the bootstrap tests, the panel, and the Java layer the
|
||||
# binary EMBEDS (bootstrap_asset.go ships the plugin sources, so a tag whose plugins do
|
||||
# not compile turns every install of that release into a failed bootstrap).
|
||||
gates:
|
||||
uses: ./.github/workflows/ci.yml
|
||||
|
||||
build:
|
||||
needs: gates
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
|
||||
# A tag that ships red is worse than a tag that fails to ship.
|
||||
- run: go vet ./...
|
||||
- run: go test ./...
|
||||
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4.4.0
|
||||
|
||||
# Both architectures, because bootstrap's default release channel DOWNLOADS these
|
||||
# rather than compiling on the target host — an arm64 host with no asset silently
|
||||
# falls back to a slow source build. Neither stage is emulated: the Dockerfile pins
|
||||
# both build stages to $BUILDPLATFORM and the Go stage cross-compiles via TARGETARCH,
|
||||
# so the second architecture costs about a minute.
|
||||
- uses: docker/setup-buildx-action@v3
|
||||
- uses: docker/setup-buildx-action@8d2750c68a42422c14e847fe6c8ac0403b4cbd6f # v3.12.0
|
||||
|
||||
- name: Build the stamped binaries
|
||||
run: |
|
||||
@@ -75,8 +86,67 @@ jobs:
|
||||
file ./felis-linux-arm64 | grep -q 'ARM aarch64' \
|
||||
|| { echo "felis-linux-arm64 is not an arm64 ELF — TARGETARCH did not reach the go build"; exit 1; }
|
||||
|
||||
# --verify-tag refuses to invent a release for a tag that is not pushed. The upload
|
||||
# fallback makes a re-run converge rather than failing on an existing release.
|
||||
- name: Checksum the binaries
|
||||
run: sha256sum felis-linux-amd64 felis-linux-arm64 | tee SHA256SUMS
|
||||
|
||||
# A CycloneDX SBOM per binary: the Go modules (and versions) linked into it, read
|
||||
# from the build info the linker embeds.
|
||||
- uses: anchore/sbom-action@e22c389904149dbc22b58101806040fa8d37a610 # v0.24.0
|
||||
with:
|
||||
file: felis-linux-amd64
|
||||
format: cyclonedx-json
|
||||
output-file: felis-linux-amd64.cdx.json
|
||||
upload-artifact: false
|
||||
upload-release-assets: false
|
||||
- uses: anchore/sbom-action@e22c389904149dbc22b58101806040fa8d37a610 # v0.24.0
|
||||
with:
|
||||
file: felis-linux-arm64
|
||||
format: cyclonedx-json
|
||||
output-file: felis-linux-arm64.cdx.json
|
||||
upload-artifact: false
|
||||
upload-release-assets: false
|
||||
|
||||
- uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4.6.2
|
||||
with:
|
||||
name: release-assets
|
||||
path: |
|
||||
felis-linux-amd64
|
||||
felis-linux-arm64
|
||||
felis-linux-amd64.cdx.json
|
||||
felis-linux-arm64.cdx.json
|
||||
SHA256SUMS
|
||||
if-no-files-found: error
|
||||
retention-days: 7
|
||||
|
||||
publish:
|
||||
needs: build
|
||||
runs-on: ubuntu-latest
|
||||
permissions:
|
||||
contents: write # gh release create/upload
|
||||
id-token: write # the Sigstore certificate behind the provenance attestation
|
||||
attestations: write
|
||||
steps:
|
||||
- uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4.3.0
|
||||
with:
|
||||
name: release-assets
|
||||
|
||||
# The artifact store sits between the two jobs, so check the handover too.
|
||||
- run: sha256sum -c SHA256SUMS
|
||||
|
||||
# Signed SLSA provenance: which workflow run, commit and repository produced each
|
||||
# binary. Check one with `gh attestation verify felis-linux-amd64 --repo FelisMC/Felis`.
|
||||
# GitHub only stores attestations for private repositories on Enterprise Cloud, and a
|
||||
# failure here would block the release, so a private repository skips the step and
|
||||
# relies on SHA256SUMS alone.
|
||||
- name: Attest build provenance
|
||||
if: ${{ !github.event.repository.private }}
|
||||
uses: actions/attest-build-provenance@977bb373ede98d70efdf65b84cb5f73e068dcc2a # v3.0.0
|
||||
with:
|
||||
subject-path: |
|
||||
felis-linux-amd64
|
||||
felis-linux-arm64
|
||||
|
||||
# --verify-tag refuses to invent a release for a tag that is not pushed.
|
||||
#
|
||||
# The prerelease flag has to be passed explicitly: the trigger glob is v*, so v1.2.3-rc1
|
||||
# lands here too, and gh does not read semver out of the tag name. Published as a full
|
||||
@@ -84,13 +154,33 @@ jobs:
|
||||
# channel installs from and `felis update` polls — so every fresh install would get the
|
||||
# RC binary and every deployed felis-api would error on the felis component until a
|
||||
# stable tag was cut. Flagged, GitHub keeps latest pointing at the last stable release.
|
||||
#
|
||||
# A re-run (the release already exists) uploads only what is missing and never
|
||||
# replaces a published asset: hosts may already have installed it, and their
|
||||
# SHA256SUMS check would start failing against a swapped file. An asset that is
|
||||
# there with different bytes stops the job; cut a new tag instead.
|
||||
- name: Publish the release
|
||||
env:
|
||||
GH_TOKEN: ${{ github.token }}
|
||||
GH_REPO: ${{ github.repository }}
|
||||
run: |
|
||||
assets="felis-linux-amd64 felis-linux-arm64 felis-linux-amd64.cdx.json felis-linux-arm64.cdx.json SHA256SUMS"
|
||||
flags=""
|
||||
case "$GITHUB_REF_NAME" in *-*) flags="--prerelease" ;; esac
|
||||
gh release create "$GITHUB_REF_NAME" --verify-tag --generate-notes $flags \
|
||||
./felis-linux-amd64 ./felis-linux-arm64 \
|
||||
|| gh release upload "$GITHUB_REF_NAME" \
|
||||
./felis-linux-amd64 ./felis-linux-arm64 --clobber
|
||||
if ! gh release view "$GITHUB_REF_NAME" >/dev/null 2>&1; then
|
||||
# shellcheck disable=SC2086 # word-splitting the list is the point
|
||||
gh release create "$GITHUB_REF_NAME" --verify-tag --generate-notes $flags $assets
|
||||
exit 0
|
||||
fi
|
||||
# The REST payload's per-asset "digest" is GitHub's own sha256 of the stored file.
|
||||
published="$(gh api "repos/${GH_REPO}/releases/tags/${GITHUB_REF_NAME}" --jq '.assets[] | "\(.name) \(.digest)"')"
|
||||
for a in $assets; do
|
||||
have="$(printf '%s\n' "$published" | awk -v n="$a" '$1 == n { print $2 }')"
|
||||
want="sha256:$(sha256sum < "$a" | cut -d' ' -f1)"
|
||||
if [ -z "$have" ]; then
|
||||
gh release upload "$GITHUB_REF_NAME" "$a"
|
||||
elif [ "$have" != "$want" ]; then
|
||||
echo "::error::$a is already published with $have; this run built $want. Published assets are never replaced."
|
||||
exit 1
|
||||
fi
|
||||
done
|
||||
+2
-3
@@ -41,9 +41,8 @@ plugins/*/bin/
|
||||
|
||||
# ---- Local agent / loop state ----
|
||||
.claude/
|
||||
# Autohand-generated agent guide — kept on disk for local tooling, never tracked.
|
||||
# It rode in via 0c1cc59, claims precedence over CLAUDE.md, and tells agents to
|
||||
# run `go fmt` (destructive on this CRLF working tree).
|
||||
# Generated tooling guide, kept on disk for local use and never tracked. Its advice
|
||||
# to run `go fmt` is destructive on this CRLF working tree.
|
||||
AGENTS.md
|
||||
|
||||
# ---- Internal planning & design docs (excluded from the public remote per
|
||||
|
||||
+1013
File diff suppressed because it is too large.
Load diff
@@ -67,6 +67,20 @@ go test ./internal/api
|
||||
go test ./cmd/felis
|
||||
```
|
||||
|
||||
The hermetic suites run against in-memory fakes; the business stores' SQL is
|
||||
verified separately against a real Postgres, on a throwaway database whose name
|
||||
must contain `pgint` (the harness drops and recreates its schema and replays the
|
||||
embedded migrations):
|
||||
|
||||
```bash
|
||||
FELIS_TEST_PG_URL='postgres://felis:***@127.0.0.1:5432/felis_pgint?sslmode=disable' \
|
||||
go test -tags pgint ./internal/pgint/ -v
|
||||
```
|
||||
|
||||
Run it after touching anything under `internal/api/pgrepo.go`, `internal/submit`,
|
||||
or `internal/build` that speaks SQL: the fakes encode the contract, and this
|
||||
suite exists to catch the drift between the fakes and the real queries.
|
||||
|
||||
Build the CLI:
|
||||
|
||||
```bash
|
||||
|
||||
+7
-3
@@ -21,14 +21,18 @@
|
||||
# minutes. The FINAL stage is deliberately NOT pinned — it must stay on the target platform
|
||||
# or the published arm64 image would carry amd64 layers. It contains only COPY, which
|
||||
# BuildKit performs itself, so it needs no QEMU either; adding a RUN there would.
|
||||
FROM --platform=$BUILDPLATFORM node:22-bookworm AS panel
|
||||
#
|
||||
# Every base image here and in deploy/{limbo,lobby,paper} is pinned by digest, so a rebuild
|
||||
# of one release uses the same bytes; .github/dependabot.yml proposes the bumps (tag and
|
||||
# digest together).
|
||||
FROM --platform=$BUILDPLATFORM node:22-bookworm@sha256:363e1587494626837fa7f9a23bdb453d13b0ff3c67c705c2805cfc69c2d2fad7 AS panel
|
||||
WORKDIR /panel
|
||||
COPY panel/package*.json ./
|
||||
RUN npm ci
|
||||
COPY panel/ ./
|
||||
RUN npm run build
|
||||
|
||||
FROM --platform=$BUILDPLATFORM golang:1.26 AS build
|
||||
FROM --platform=$BUILDPLATFORM golang:1.27@sha256:3680233e3204827fbdc66088528ae6d4b3d034f51d03a99d454f6de034888244 AS build
|
||||
WORKDIR /src
|
||||
ARG TARGETOS=linux
|
||||
ARG TARGETARCH
|
||||
@@ -55,7 +59,7 @@ ARG FELIS_VERSION=dev
|
||||
RUN CGO_ENABLED=0 GOOS="$TARGETOS" GOARCH="${TARGETARCH:-$(go env GOARCH)}" \
|
||||
go build -trimpath -ldflags="-s -w -X main.version=${FELIS_VERSION}" -o /out/felis ./cmd/felis
|
||||
|
||||
FROM gcr.io/distroless/static-debian12:nonroot
|
||||
FROM gcr.io/distroless/static-debian12:nonroot@sha256:afa5c872c891853ca7fcf1f12c3edb23f7eeef36189728842dd51042ff57f7ab
|
||||
ENV PATH=/usr/local/bin:/usr/bin:/bin
|
||||
COPY --chmod=0755 --from=build /out/felis /usr/local/bin/felis
|
||||
# distroless "nonroot" is uid 65532; the rendered PodSecurityContext pins
|
||||
|
||||
@@ -17,10 +17,11 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
|
||||
|
||||
- **即开即玩**:玩家尝试连接时自动唤醒服务器,空闲后自动休眠,像游戏主机一样省资源。
|
||||
- **Web 控制面板**:浏览器中查看服务器状态、在线玩家与资源用量,管理备份与恢复。
|
||||
- **自动备份与恢复**:定时将世界打包存档,支持从任意备份点一键回滚。
|
||||
- **智慧回收**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间。
|
||||
- **备份与恢复**:一键把整服数据(世界、配置、插件/模组,即整个 /data 卷)打包进集群内的归档库,支持从任意备份点回滚;默认安装就已启用(归档 PVC 与路径由安装器一并生成)。
|
||||
- **控制面数据库备份**:账号、服务器归属、配额与存档索引所在的数据库每天自动备份,每次升级迁移前先快照,出错可用 `felis db restore` 整库原子回滚;面板「维护与备份」页显示备份是否新鲜(见 [故障排查 §16](docs/troubleshooting.md))。
|
||||
- **智慧回收(可选开启)**:超过 15 天无人游玩的世界自动备份后删除,释放磁盘空间;安装时设置 `FELIS_WORLDS_HOST_PATH`(k3s 默认 `/var/lib/rancher/k3s/storage`)即启用每日回收,不设置则不删任何世界。
|
||||
- **多核心支持**:兼容 Paper、Fabric、Forge、NeoForge,经由 Velocity 代理统一入口。
|
||||
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建并部署。
|
||||
- **模组自助提交**:玩家自行上传模组包,服主审批通过后自动构建;构建产物进入镜像白名单,可直接选用为服务器镜像完成部署。
|
||||
- **Passkey 登录**:支持指纹、面容、硬件密钥等无密码认证方式。
|
||||
- **零信任安全**:面板流量由 Cloudflare Access 保护,集群内 API 不暴露到公网。
|
||||
|
||||
@@ -29,7 +30,7 @@ A Kubernetes-driven Minecraft server hosting platform — one command to deploy,
|
||||
在准备好的 Linux 主机上执行:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
```
|
||||
|
||||
脚本将自动安装 K3s、部署控制平面并启动设置向导。完成后浏览器访问已配置的域名进入控制面板即可使用。
|
||||
@@ -41,7 +42,7 @@ curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/de
|
||||
> export FELIS_GITHUB_TOKEN=<对本仓库有读权限的 token>
|
||||
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
|
||||
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
|
||||
> https://api.github.com/repos/MliroLirrorsIngenuity/Felis/contents/deploy/bootstrap.sh \
|
||||
> https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
|
||||
> | sudo -E bash
|
||||
> ```
|
||||
>
|
||||
|
||||
+26
-4
@@ -17,10 +17,11 @@ Table of Contents
|
||||
|
||||
- **Wake on Join**: Servers start automatically when a player connects, and stop when idle — like hibernate for your server.
|
||||
- **Web Dashboard**: Monitor server status, online players, and resource usage from your browser, with backup and restore management.
|
||||
- **Auto Backup & Restore**: Scheduled world backups with one-click rollback from any backup point.
|
||||
- **World Reaper**: Worlds idle for more than 15 days are automatically backed up and removed to free disk space.
|
||||
- **Backup & Restore**: One-click snapshots of a server's whole data volume (worlds, config, plugins/mods — the entire /data volume) into the cluster's archive store, with rollback from any backup point — enabled by default (the installer renders the archive PVC and its path).
|
||||
- **Control-plane database backups**: The database holding accounts, server ownership, quotas and the archive index is backed up daily and snapshotted before every upgrade migrates it; `felis db restore` rolls it back atomically, and the panel's Maintenance & Backups page shows whether the newest backup is fresh (see [troubleshooting §16](docs/troubleshooting.md)).
|
||||
- **World Reaper** (opt in): Worlds idle for more than 15 days are automatically backed up and removed to free disk space. Enable it by setting `FELIS_WORLDS_HOST_PATH` at install time (on k3s: `/var/lib/rancher/k3s/storage`); without it, no world is ever deleted.
|
||||
- **Multi-core Support**: Compatible with Paper, Fabric, Forge, and NeoForge, federated behind a Velocity proxy.
|
||||
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers automatic build and deployment.
|
||||
- **Modpack Submission**: Players submit custom modpacks; admin approval triggers an automatic build, and the result is whitelisted as a server image you can select to deploy.
|
||||
- **Passkey Login**: Passwordless authentication via fingerprint, face recognition, or hardware security keys.
|
||||
- **Zero Trust Security**: Panel traffic protected by Cloudflare Access; the internal API is never exposed to the internet.
|
||||
|
||||
@@ -29,11 +30,32 @@ Table of Contents
|
||||
On a prepared Linux host, run:
|
||||
|
||||
```bash
|
||||
curl -fsSL https://raw.githubusercontent.com/MliroLirrorsIngenuity/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/main/deploy/bootstrap.sh | sudo bash
|
||||
```
|
||||
|
||||
The script installs K3s, deploys the control plane, and launches a setup wizard. Once done, open your browser at the configured domain.
|
||||
|
||||
> **This repository is currently private**, so the command above returns 404. Use the
|
||||
> credentialed form instead; the installer itself needs the same token to resolve and
|
||||
> download the release, so pass it through with `sudo -E`:
|
||||
>
|
||||
> ```bash
|
||||
> export FELIS_GITHUB_TOKEN=<a token with read access to this repository>
|
||||
> printf 'header = "Authorization: Bearer %s"\n' "$FELIS_GITHUB_TOKEN" \
|
||||
> | curl -fsSL --config - -H "Accept: application/vnd.github.raw" \
|
||||
> https://api.github.com/repos/FelisMC/Felis/contents/deploy/bootstrap.sh \
|
||||
> | sudo -E bash
|
||||
> ```
|
||||
>
|
||||
> The token reaches `curl --config -` over stdin instead of the command line: argv is
|
||||
> readable by any local user via `/proc`, and that is exactly why the installer's
|
||||
> internal `github_api` uses the same form.
|
||||
|
||||
Rerunning this command is also how you upgrade felis-api to a newer version (`felis setup`
|
||||
cannot — it uses the binary already installed on the host). The rerun keeps the installed
|
||||
root domain but **not** the channel: if this host follows main, also
|
||||
`export FELIS_VERSION_BOOTSTRAP=dev`.
|
||||
|
||||
## Build from Source
|
||||
|
||||
Felis is built with Go and Node.js:
|
||||
|
||||
+7
-5
@@ -13,8 +13,8 @@ var bootstrapScript string
|
||||
//go:embed deploy/crd/*.yaml
|
||||
var bootstrapAssets embed.FS
|
||||
|
||||
// gameStackAssets carries everything deploy/bootstrap.sh needs to build the two
|
||||
// always-on game images (login limbo + lobby) and the Velocity plugin, for the TUI
|
||||
// gameStackAssets carries everything deploy/bootstrap.sh needs to build the three
|
||||
// game images (login limbo, lobby, plain Paper) and the Velocity plugin, for the TUI
|
||||
// install path — which pipes the embedded bootstrap.sh into bash and therefore has
|
||||
// NO source checkout on disk to build from.
|
||||
//
|
||||
@@ -24,8 +24,10 @@ var bootstrapAssets embed.FS
|
||||
// otherwise be baked into every felis binary. Keep them explicit — add a source
|
||||
// directory here, never a parent.
|
||||
//
|
||||
//go:embed deploy/game-stack.lock
|
||||
//go:embed deploy/limbo/Dockerfile deploy/limbo/entrypoint.sh
|
||||
//go:embed deploy/lobby/Dockerfile deploy/lobby/entrypoint.sh
|
||||
//go:embed deploy/paper/Dockerfile deploy/paper/entrypoint.sh
|
||||
//go:embed plugins/limbo/build.gradle plugins/limbo/settings.gradle plugins/limbo/src
|
||||
//go:embed plugins/paper/build.gradle plugins/paper/settings.gradle plugins/paper/src
|
||||
//go:embed plugins/velocity/build.gradle plugins/velocity/settings.gradle plugins/velocity/src
|
||||
@@ -46,9 +48,9 @@ func GameStackTar(w io.Writer) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
// Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by both
|
||||
// Dockerfiles precisely because the +x bit does not survive a Windows checkout,
|
||||
// so nothing here needs to be executable.
|
||||
// Mode 0644 for everything: entrypoint.sh is invoked as `sh <file>` by all
|
||||
// three Dockerfiles precisely because the +x bit does not survive a Windows
|
||||
// checkout, so nothing here needs to be executable.
|
||||
if err := tw.WriteHeader(&tar.Header{
|
||||
Name: path,
|
||||
Mode: 0o644,
|
||||
|
||||
@@ -1,6 +1,9 @@
|
||||
package felis
|
||||
|
||||
import (
|
||||
"io/fs"
|
||||
"os"
|
||||
"regexp"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
@@ -64,6 +67,37 @@ func TestLobbyLuckPermsWiringIsConsistent(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The Paper jar digest rides the same cross-file contract as LuckPerms above: bootstrap.sh
|
||||
// resolves "url sha256" out of Fill's content-addressed download URL and passes the digest
|
||||
// as a build-arg the Dockerfile must require and verify. docker only WARNS about an unknown
|
||||
// --build-arg, so a renamed arg would surface as a required-arg failure on a real host
|
||||
// mid-install — this test is the only compile step the pairing gets.
|
||||
//
|
||||
// Both images pull the same jar from the same URL, so both have to check it: a gate on one
|
||||
// of them leaves the other booting on whatever bytes happened to arrive.
|
||||
func TestPaperJarDigestWiringIsConsistent(t *testing.T) {
|
||||
const arg = "PAPER_JAR_SHA256"
|
||||
if n := strings.Count(BootstrapScript(), "--build-arg "+arg+"="); n < 2 {
|
||||
t.Errorf("bootstrap.sh passes --build-arg %s %d time(s); the lobby and the "+
|
||||
"plain-Paper build each need it", arg, n)
|
||||
}
|
||||
for _, name := range []string{"deploy/lobby/Dockerfile", "deploy/paper/Dockerfile"} {
|
||||
dockerfile := readGameStackFile(t, name)
|
||||
if !strings.Contains(dockerfile, "ARG "+arg) {
|
||||
t.Errorf("%s declares no ARG %s", name, arg)
|
||||
}
|
||||
if !strings.Contains(dockerfile, `if [ -z "${PAPER_JAR_SHA256:-}" ]`) {
|
||||
t.Errorf("%s does not fail the build when %s is unset", name, arg)
|
||||
}
|
||||
// Requiring the arg is not the same as spending it, and which file gets hashed
|
||||
// matters as much as the command: a `sha256sum -c` over some other download
|
||||
// would satisfy a bare substring check while paper.jar still arrives unchecked.
|
||||
if !strings.Contains(dockerfile, `echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c`) {
|
||||
t.Errorf("%s never verifies /paper/paper.jar against %s", name, arg)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A 1.8 client joining a protocol-47 backend dies on the first chunk unless ViaVersion's
|
||||
// serverside block-connection tracking is off: under modern forwarding the Velocity injector
|
||||
// reports 1.13 as the lowest supported protocol, ConnectionData.init() returns early on that,
|
||||
@@ -104,6 +138,181 @@ func TestBootstrapPinsViaBlockConnectionsOff(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The embed list and the images bootstrap.sh builds are two lists nobody reconciles.
|
||||
// deploy/paper shipped an image build without ever being added to gameStackAssets, and
|
||||
// nothing said so: a checkout on disk satisfies the build either way, and the tar is
|
||||
// only the build context on the path that has no checkout — `curl | bash`, where the
|
||||
// third `docker build -f` then names a file that was never unpacked. So derive the
|
||||
// inputs from the script and from each Dockerfile's own COPY lines instead of restating
|
||||
// them here; a fourth image inherits the check for free.
|
||||
func TestGameStackTarCarriesEveryBuildInput(t *testing.T) {
|
||||
// Matches the path only when GAME_STACK_DIR is followed by one, which skips the
|
||||
// build-context arguments (`"$GAME_STACK_DIR"`, `"${GAME_STACK_DIR}:/src:z"`) and
|
||||
// the glob for gradle's output, none of which are inputs this tar has to carry.
|
||||
found := regexp.MustCompile(`\$\{GAME_STACK_DIR\}/(\S+?)"`).FindAllStringSubmatch(BootstrapScript(), -1)
|
||||
var paths []string
|
||||
seen := map[string]bool{}
|
||||
for _, m := range found {
|
||||
if !seen[m[1]] {
|
||||
seen[m[1]] = true
|
||||
paths = append(paths, m[1])
|
||||
}
|
||||
}
|
||||
// Guards the regex itself: a rewrite of how bootstrap.sh spells the build context
|
||||
// would otherwise turn this test into an unconditional pass. It has to come before
|
||||
// the loop — a missing file in there is fatal, and a floor placed after it would
|
||||
// never be reached to say that the regex, not the tar, is what went wrong.
|
||||
if len(paths) < 3 {
|
||||
t.Fatalf("only %d game-stack path(s) resolved out of bootstrap.sh; the limbo, "+
|
||||
"lobby and paper Dockerfiles are all built from ${GAME_STACK_DIR}", len(paths))
|
||||
}
|
||||
for _, path := range paths {
|
||||
requireEmbedded(t, path)
|
||||
// A Dockerfile that arrives without the files it COPYs fails just as late and
|
||||
// just as far from here; the deploy/paper gap was missing its entrypoint too.
|
||||
for _, src := range copySources(t, path) {
|
||||
requireEmbedded(t, src)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// copySources lists the build-context paths a Dockerfile COPYs in, skipping the
|
||||
// --from=<stage> copies, whose sources are produced by an earlier stage rather than
|
||||
// unpacked from the tar.
|
||||
func copySources(t *testing.T, dockerfile string) []string {
|
||||
t.Helper()
|
||||
var out []string
|
||||
// Continuations are joined first: a COPY split across lines would otherwise be two
|
||||
// fragments, neither of them starting with COPY followed by a source, and its
|
||||
// source would slip past unchecked.
|
||||
body := strings.ReplaceAll(readGameStackFile(t, dockerfile), "\\\n", " ")
|
||||
for line := range strings.SplitSeq(body, "\n") {
|
||||
f := strings.Fields(line)
|
||||
if len(f) < 2 || f[0] != "COPY" || strings.HasPrefix(f[1], "--") {
|
||||
continue
|
||||
}
|
||||
out = append(out, strings.TrimSuffix(f[1], "/"))
|
||||
}
|
||||
return out
|
||||
}
|
||||
|
||||
// fs.Stat rather than ReadFile: half of these are directories (`COPY plugins/shared/`),
|
||||
// and embed.FS answers for those too.
|
||||
func requireEmbedded(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
if _, err := fs.Stat(gameStackAssets, path); err != nil {
|
||||
t.Errorf("%s is a game-stack build input but is not in gameStackAssets; an "+
|
||||
"install with no source checkout dies on it: %v", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
// The lock file is the install's only source of upstream builds, and bootstrap.sh reads it
|
||||
// with a strict KEY=value parser that dies on anything unexpected, so a malformed lock is a
|
||||
// failed install on every host. Check the shipped copy the same way here.
|
||||
func TestGameStackLockIsComplete(t *testing.T) {
|
||||
lock := map[string]string{}
|
||||
for line := range strings.SplitSeq(readGameStackFile(t, "deploy/game-stack.lock"), "\n") {
|
||||
if line == "" || strings.HasPrefix(line, "#") {
|
||||
continue
|
||||
}
|
||||
k, v, ok := strings.Cut(line, "=")
|
||||
if !ok {
|
||||
t.Fatalf("not a KEY=value line: %q", line)
|
||||
}
|
||||
lock[k] = v
|
||||
}
|
||||
m := regexp.MustCompile(`GAME_STACK_LOCK_KEYS="([^"]*)"`).FindStringSubmatch(BootstrapScript())
|
||||
if m == nil {
|
||||
t.Fatal("bootstrap.sh no longer declares GAME_STACK_LOCK_KEYS")
|
||||
}
|
||||
keys := strings.Fields(m[1])
|
||||
sha := regexp.MustCompile(`^[0-9a-f]{64}$`)
|
||||
for _, k := range keys {
|
||||
v, ok := lock[k]
|
||||
if !ok || v == "" {
|
||||
t.Errorf("game-stack.lock does not set %s", k)
|
||||
continue
|
||||
}
|
||||
if strings.HasSuffix(k, "_SHA256") && !sha.MatchString(v) {
|
||||
t.Errorf("%s=%q is not a lowercase sha256", k, v)
|
||||
}
|
||||
// A moving URL pins nothing: the digest check would start failing the day
|
||||
// upstream publishes the next build.
|
||||
if strings.HasSuffix(k, "_URL") && strings.Contains(v, "lastSuccessfulBuild") {
|
||||
t.Errorf("%s names a moving build: %s", k, v)
|
||||
}
|
||||
}
|
||||
for k := range lock {
|
||||
if !strings.Contains(" "+m[1]+" ", " "+k+" ") {
|
||||
t.Errorf("game-stack.lock sets %s, which bootstrap.sh refuses as an unknown key", k)
|
||||
}
|
||||
}
|
||||
// Fill's URLs are content-addressed; a lock whose digest disagrees with its own URL
|
||||
// was edited by hand and half-way.
|
||||
for _, name := range []string{"PAPER", "VELOCITY"} {
|
||||
if !strings.Contains(lock[name+"_JAR_URL"], "/objects/"+lock[name+"_JAR_SHA256"]+"/") {
|
||||
t.Errorf("%s_JAR_SHA256 is not the digest in %s_JAR_URL", name, name)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(lock["LIMBO_JAR_URL"], "-"+lock["MC_VERSION"]+".jar") {
|
||||
t.Errorf("LIMBO_JAR_URL %s is not a Minecraft %s build", lock["LIMBO_JAR_URL"], lock["MC_VERSION"])
|
||||
}
|
||||
if !strings.Contains(lock["PAPER_JAR_URL"], "/paper-"+lock["MC_VERSION"]+"-") {
|
||||
t.Errorf("PAPER_JAR_URL %s is not a Minecraft %s build; the lobby would not speak the login gate's protocol", lock["PAPER_JAR_URL"], lock["MC_VERSION"])
|
||||
}
|
||||
}
|
||||
|
||||
// Each downloaded jar's digest is a build-arg bootstrap.sh passes and the Dockerfile must
|
||||
// both require and spend on the file it downloaded; docker only warns about an unknown
|
||||
// --build-arg, so a renamed arg would ship an unchecked jar.
|
||||
func TestGameStackDigestsReachTheImageBuilds(t *testing.T) {
|
||||
script := BootstrapScript()
|
||||
for _, c := range []struct{ dockerfile, arg, path string }{
|
||||
{"deploy/limbo/Dockerfile", "LIMBO_JAR_SHA256", "/limbo/Limbo.jar"},
|
||||
{"deploy/limbo/Dockerfile", "LIMBO_SCHEM_SHA256", "/limbo/spawn.schem"},
|
||||
{"deploy/lobby/Dockerfile", "LUCKPERMS_JAR_SHA256", "/paper/plugins/LuckPerms.jar"},
|
||||
} {
|
||||
if !strings.Contains(script, "--build-arg "+c.arg+"=\"$"+c.arg+"\"") {
|
||||
t.Errorf("bootstrap.sh never passes --build-arg %s", c.arg)
|
||||
}
|
||||
dockerfile := readGameStackFile(t, c.dockerfile)
|
||||
if !strings.Contains(dockerfile, "ARG "+c.arg) {
|
||||
t.Errorf("%s declares no ARG %s", c.dockerfile, c.arg)
|
||||
}
|
||||
if !strings.Contains(dockerfile, `echo "$`+c.arg+` `+c.path+`" | sha256sum -c`) {
|
||||
t.Errorf("%s never verifies %s against %s", c.dockerfile, c.path, c.arg)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A base image named by tag alone is whatever the tag points at on build day.
|
||||
func TestDockerfileBaseImagesArePinnedByDigest(t *testing.T) {
|
||||
root, err := os.ReadFile("Dockerfile")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
files := map[string]string{"Dockerfile": string(root)}
|
||||
for _, name := range []string{"deploy/limbo/Dockerfile", "deploy/lobby/Dockerfile", "deploy/paper/Dockerfile"} {
|
||||
files[name] = readGameStackFile(t, name)
|
||||
}
|
||||
pinned := regexp.MustCompile(`^FROM (--platform=\S+ )?\S+:\S+@sha256:[0-9a-f]{64}( AS \S+)?$`)
|
||||
for name, body := range files {
|
||||
n := 0
|
||||
for line := range strings.SplitSeq(body, "\n") {
|
||||
if !strings.HasPrefix(line, "FROM ") {
|
||||
continue
|
||||
}
|
||||
n++
|
||||
if !pinned.MatchString(line) {
|
||||
t.Errorf("%s: %q is not pinned by digest", name, line)
|
||||
}
|
||||
}
|
||||
if n == 0 {
|
||||
t.Errorf("%s has no FROM line", name)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func readGameStackFile(t *testing.T, name string) string {
|
||||
t.Helper()
|
||||
b, err := gameStackAssets.ReadFile(name)
|
||||
|
||||
+264
-24
@@ -5,10 +5,12 @@ import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"os"
|
||||
"regexp"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/api"
|
||||
@@ -17,10 +19,15 @@ import (
|
||||
"felis.lolicon.best/internal/build"
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/fileedit"
|
||||
"felis.lolicon.best/internal/imagepin"
|
||||
"felis.lolicon.best/internal/mail"
|
||||
"felis.lolicon.best/internal/metrics"
|
||||
"felis.lolicon.best/internal/naming"
|
||||
"felis.lolicon.best/internal/panel"
|
||||
"felis.lolicon.best/internal/passkey"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"felis.lolicon.best/internal/reaper"
|
||||
"felis.lolicon.best/internal/registryprune"
|
||||
"felis.lolicon.best/internal/restore"
|
||||
"felis.lolicon.best/internal/store"
|
||||
"felis.lolicon.best/internal/submit"
|
||||
@@ -108,6 +115,8 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
return 1
|
||||
}
|
||||
|
||||
metrics.SetBuildInfo("api", resolvedVersion())
|
||||
|
||||
token := os.Getenv("FELIS_SERVICE_TOKEN")
|
||||
if token == "" {
|
||||
fmt.Fprintln(stderr, "felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers")
|
||||
@@ -142,11 +151,17 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
// build namespace and pushes to the internal registry. The build Pod never
|
||||
// holds DB credentials — felis-api owns the PG store and admits scanned
|
||||
// images, so the Builder is constructed here with both bindings.
|
||||
buildCfg := buildConfig(cfg)
|
||||
// The fetch initContainer runs THIS image's fetch-context entrypoint, so the
|
||||
// build config carries the api's own image (the platform sets FELIS_IMAGE).
|
||||
buildCfg.FelisImage = os.Getenv("FELIS_IMAGE")
|
||||
buildJobs := build.NewK8sJobs(cl, buildCfg)
|
||||
builder := &build.Builder{
|
||||
Store: build.NewPGStore(drv.DB()),
|
||||
Jobs: build.NewK8sJobs(cl, buildConfig(cfg)),
|
||||
Config: buildConfig(cfg),
|
||||
Jobs: buildJobs,
|
||||
Config: buildCfg,
|
||||
}
|
||||
go probeBuildUserNamespaces(ctx, buildJobs, buildCfg, stderr)
|
||||
|
||||
// User-modpack approval lane (user-directed extension over §16; see
|
||||
// internal/submit). An ordinary user may only SUBMIT a
|
||||
@@ -159,14 +174,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
// The blob upload transport is selected by the shape of user_uploads_context —
|
||||
// the two backends the setup wizard chooses between. A local path wires
|
||||
// LocalContextStore (the mounted uploads PVC); an s3:// base wires
|
||||
// S3ContextStore when its credentials resolve. Either way the store's target is
|
||||
// derived from the SAME config field the context ref uses, so the blob lands
|
||||
// exactly where Kaniko's --context points. Anything else — or an s3:// base with
|
||||
// no credentials configured — leaves Blobs nil so POST
|
||||
// S3ContextStore when its credentials resolve. Anything else — or an s3:// base
|
||||
// with no credentials configured — leaves Blobs nil so POST
|
||||
// /me/submissions/{id}/context returns 503, honest like the restore executor
|
||||
// when its PVC is not supplied. (Letting the sandboxed Kaniko build Pod READ the
|
||||
// context — PVC mount for local, creds+egress for S3 — is a separate deployment
|
||||
// integration.)
|
||||
// when its PVC is not supplied.
|
||||
//
|
||||
// Reading the blob back is the API's job, not Kaniko's: the build Pod runs in
|
||||
// another namespace and can neither mount the uploads PVC (a PVC does not cross
|
||||
// namespaces) nor hold object-store credentials, so ContextBaseURL makes the
|
||||
// derived context ref an internal-face URL that the build Job's fetch
|
||||
// initContainer streams (cmd/felis fetch-context). The platform renders this
|
||||
// address into the api Deployment (felis API base URL env); the fallback keeps
|
||||
// a hand-rolled deployment working under the platform's default control
|
||||
// namespace.
|
||||
contextBase := cfg.Registry.UserUploadsContext
|
||||
var blobs submit.Blobs
|
||||
switch {
|
||||
@@ -186,11 +206,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase)
|
||||
}
|
||||
submissions := &submit.Manager{
|
||||
Store: submit.NewPGStore(drv.DB()),
|
||||
Builds: builder,
|
||||
Registry: cfg.Registry.URL,
|
||||
ContextStore: contextBase,
|
||||
Blobs: blobs,
|
||||
Store: submit.NewPGStore(drv.DB()),
|
||||
Builds: builder,
|
||||
Registry: cfg.Registry.URL,
|
||||
ContextStore: contextBase,
|
||||
ContextBaseURL: internalAPIBaseURL(),
|
||||
Blobs: blobs,
|
||||
}
|
||||
if v := cfg.Registry.UserUploadsMaxBytes; v != "" {
|
||||
if n, err := parseByteSize(v); err != nil || n <= 0 {
|
||||
fmt.Fprintf(stderr, "felis api: [registry] user_uploads_max_bytes %q is not a positive size such as 4Gi; keeping the default\n", v)
|
||||
} else {
|
||||
submissions.MaxStoredBytesTotal = n
|
||||
}
|
||||
}
|
||||
|
||||
// Restore subsystem (spec §7): the weak-SA restore Job mounts the target
|
||||
@@ -247,9 +275,19 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
// agree on what local auth knows.
|
||||
repo := api.NewPGRepo(drv.DB())
|
||||
|
||||
// The owner's on-demand backup levers come from [archive], the same keys the
|
||||
// backup Job and the reaper read. A malformed key leaves the defaults in
|
||||
// place here; the reaper Job fails on it and names it.
|
||||
rcfg, err := reaperConfig(cfg)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis api: %v; using the default backup limits\n", err)
|
||||
rcfg = reaper.DefaultConfig()
|
||||
}
|
||||
|
||||
cluster := api.NewK8sCluster(cl, cfg.K8s.Namespace)
|
||||
a := &api.API{
|
||||
Repo: repo,
|
||||
Cluster: api.NewK8sCluster(cl, cfg.K8s.Namespace),
|
||||
Cluster: cluster,
|
||||
Console: api.NewK8sConsole(cl, cfg.K8s.Namespace),
|
||||
Logs: api.NewK8sLogStreamer(clientset, cfg.K8s.Namespace),
|
||||
// Build-log stream (spec §16) is scoped to the BUILD namespace — the same
|
||||
@@ -257,8 +295,10 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
BuildLogs: api.NewK8sBuildLogStreamer(clientset, cfg.Registry.BuildNamespace),
|
||||
Internal: api.BearerTokenAuth{Token: token},
|
||||
Builder: builder,
|
||||
Images: imagePinner(cfg.Registry.URL),
|
||||
Restorer: restorer,
|
||||
Backuper: backuper,
|
||||
JobStatus: api.NewK8sJobStatus(cl, cfg.K8s.Namespace),
|
||||
Files: files,
|
||||
Submissions: submissions,
|
||||
Mailer: mailer,
|
||||
@@ -279,23 +319,43 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
AdminHostname: cfg.Auth.AdminHostname,
|
||||
PanelHostname: cfg.Auth.PanelHostname,
|
||||
WakeCooldown: 30 * time.Second,
|
||||
// An owner may start one backup per server per manual_cooldown, and none
|
||||
// while the store is at max_local_bytes (data-durability-9).
|
||||
BackupCooldown: rcfg.ManualCooldown,
|
||||
BackupStoreCap: rcfg.MaxLocalBytes,
|
||||
// The user-modpack lane's per-user throttles: a create spaces out
|
||||
// review-queue rows, an upload spaces out (up to 1 GiB) context streams.
|
||||
// Separate keys, so the normal create→upload sequence stays immediate.
|
||||
SubmitCreateCooldown: 30 * time.Second,
|
||||
SubmitUploadCooldown: 15 * time.Second,
|
||||
// Bound concurrent console/build-log SSE streams per principal. Generous enough
|
||||
// for legitimate multi-tab / multi-server watching, while capping how many
|
||||
// upstream follow connections a single caller can tie up if their streams stall.
|
||||
MaxStreamsPerPrincipal: 16,
|
||||
// Public sign-in doors, per client address: a person signing in makes a
|
||||
// handful of calls, so 20 at once refilled at 20 a minute never bites a
|
||||
// real user and still turns a spray into a trickle. The client address
|
||||
// is the edge's header when the install names one (config.AuthConfig).
|
||||
AuthDoorLimit: api.RateLimit{Burst: 20, PerMinute: 20},
|
||||
ClientIPHeader: cfg.Auth.EffectiveClientIPHeader(),
|
||||
MailLimit: mailLimit(cfg.SMTP.MaxPerHour),
|
||||
}
|
||||
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
|
||||
|
||||
// Felis-nano: wire the multi-source hasJoined multiplexer only when third-party auth
|
||||
// sources are configured. Mojang leads as the code-owned identity anchor (正版优先);
|
||||
// config can only append namespace-rewritten third-party sources, never a trusted one,
|
||||
// so a misconfig cannot reopen the impersonation hole. No sources = a.AuthSources stays
|
||||
// nil = the endpoint 204s every login (ships off).
|
||||
if len(cfg.AuthSources) > 0 {
|
||||
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
|
||||
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
|
||||
if a.ClientIPHeader != "" {
|
||||
fmt.Fprintf(stderr, "felis api: sign-in rate limit keys on the %s header\n", a.ClientIPHeader)
|
||||
} else {
|
||||
fmt.Fprintln(stderr, "felis api: sign-in rate limit keys on the TCP peer ([auth] client_ip_header unset)")
|
||||
}
|
||||
|
||||
// Felis-nano: the multi-source hasJoined multiplexer. Mojang leads as the code-owned
|
||||
// identity anchor (正版优先); config can only append namespace-rewritten third-party
|
||||
// sources, never a trusted one, so a misconfig cannot reopen the impersonation hole.
|
||||
// Wired unconditionally: the installer points Velocity at this route whether or not any
|
||||
// [[auth_source]] is configured, so an empty list has to mean a Mojang-only relay, the
|
||||
// same as under `felis nano`. A nil list would 204 every login, premium ones included.
|
||||
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
|
||||
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
|
||||
|
||||
// Passkey (WebAuthn) verifier (spec §14, Phase 6). One relying party spans BOTH
|
||||
// web faces: the RP id is the panel hostname (console.<root>), and because that is
|
||||
// a domain suffix of the operator host (op.console.<root>), a single credential
|
||||
@@ -349,6 +409,11 @@ func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
||||
// reconciles it, but this loop converges builds nobody is polling.
|
||||
go reconcileBuilds(ctx, builder, stderr)
|
||||
|
||||
if pruner := registryPruner(cfg, builder.Store, cluster, stderr); pruner != nil {
|
||||
go pruner.Loop(ctx, registryPruneInterval)
|
||||
}
|
||||
go reapRejectedContexts(ctx, submissions, stderr)
|
||||
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
|
||||
@@ -402,9 +467,60 @@ func buildConfig(cfg *config.Config) build.Config {
|
||||
return build.Config{
|
||||
Namespace: cfg.Registry.BuildNamespace,
|
||||
RegistryURL: cfg.Registry.URL,
|
||||
// Empty overrides fall back to the registry's copies of the tools
|
||||
// (build.Tools), which felis mirror-build-tools keeps current.
|
||||
KanikoImage: cfg.Registry.KanikoImage,
|
||||
TrivyImage: cfg.Registry.TrivyImage,
|
||||
CPULimit: cfg.Registry.BuildCPULimit,
|
||||
MemLimit: cfg.Registry.BuildMemLimit,
|
||||
DiskLimit: cfg.Registry.BuildDiskLimit,
|
||||
// "auto" follows the startup probe (see probeBuildUserNamespaces).
|
||||
UserNamespaces: cfg.Registry.BuildUserNamespaces,
|
||||
UserNamespacesProbe: new(atomic.Bool),
|
||||
RuntimeClass: cfg.Registry.BuildRuntimeClass,
|
||||
MaxConcurrent: cfg.Registry.MaxConcurrentBuilds,
|
||||
TrivyDBRepository: cfg.Registry.TrivyDBRepository,
|
||||
TrivyJavaDBRepository: cfg.Registry.TrivyJavaDBRepository,
|
||||
// The submit lane's derived context URLs live here; the fetch step's
|
||||
// service token goes nowhere else.
|
||||
ContextOrigin: internalAPIBaseURL(),
|
||||
}
|
||||
}
|
||||
|
||||
// probeBuildUserNamespaces settles build_user_namespaces = "auto": one probe
|
||||
// pod with hostUsers: false tells whether this node's kernel and runtime can run
|
||||
// build pods in a user namespace. Builds submitted before it answers run without.
|
||||
func probeBuildUserNamespaces(ctx context.Context, jobs *build.K8sJobs, cfg build.Config, stderr io.Writer) {
|
||||
if mode := cfg.UserNamespaces; mode != "" && mode != build.UserNamespacesAuto {
|
||||
return
|
||||
}
|
||||
if cfg.FelisImage == "" {
|
||||
fmt.Fprintln(stderr, "felis api: FELIS_IMAGE unset — build pods run without a user namespace")
|
||||
return
|
||||
}
|
||||
ok, err := jobs.ProbeUserNamespaces(ctx, cfg.FelisImage)
|
||||
cfg.UserNamespacesProbe.Store(ok)
|
||||
switch {
|
||||
case ok:
|
||||
fmt.Fprintln(stderr, "felis api: build pods run in a user namespace (hostUsers: false)")
|
||||
case err != nil:
|
||||
fmt.Fprintf(stderr, "felis api: build pods run without a user namespace: the probe failed: %v\n", err)
|
||||
default:
|
||||
fmt.Fprintln(stderr, "felis api: build pods run without a user namespace: this node cannot start a pod with hostUsers: false")
|
||||
}
|
||||
}
|
||||
|
||||
// internalAPIBaseURL resolves the platform's internal-face base URL: the address
|
||||
// the platform rendered into this pod (felis API base URL env), or — for a
|
||||
// hand-rolled deployment that set none — the platform default control namespace,
|
||||
// the same fallback setup.go uses to hand the login gate its address.
|
||||
func internalAPIBaseURL() string {
|
||||
if base := os.Getenv(naming.EnvAPIBaseURL); base != "" {
|
||||
return base
|
||||
}
|
||||
return platform.InternalAPIBaseURL(platform.DefaultControlNamespace)
|
||||
}
|
||||
|
||||
// uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://".
|
||||
var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`)
|
||||
|
||||
@@ -507,3 +623,127 @@ func reconcileBuilds(ctx context.Context, b *build.Builder, stderr io.Writer) {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// reapRejectedContexts deletes, once an hour, the uploaded contexts of
|
||||
// submissions rejected more than submit.RejectedContextRetention ago. Without it a
|
||||
// rejected modpack keeps its bytes on the uploads store (and against its
|
||||
// submitter's budget) until an admin deletes the row.
|
||||
func reapRejectedContexts(ctx context.Context, m *submit.Manager, stderr io.Writer) {
|
||||
t := time.NewTicker(time.Hour)
|
||||
defer t.Stop()
|
||||
for {
|
||||
n, err := m.ReapRejected(ctx, submit.RejectedContextRetention)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis api: reap rejected uploads: %v\n", err)
|
||||
}
|
||||
if n > 0 {
|
||||
fmt.Fprintf(stderr, "felis api: deleted the uploaded contexts of %d rejected submission(s)\n", n)
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return
|
||||
case <-t.C:
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// registryPruneInterval spaces the registry pruner's runs. The registry-gc
|
||||
// sidecar sweeps once a day, so pruning more often only changes which sweep frees
|
||||
// a layer.
|
||||
const registryPruneInterval = 6 * time.Hour
|
||||
|
||||
// registryPruner deletes the registry manifests nothing references
|
||||
// (internal/registryprune); the registry-gc sidecar frees their layers on its next
|
||||
// sweep. It acts as the gate's prune principal, whose token the api Deployment
|
||||
// injects from felis-registry-auth. Without the token the registry only grows,
|
||||
// which is said once here.
|
||||
func registryPruner(cfg *config.Config, store imageRefStore, servers serverLister, stderr io.Writer) *registryprune.Pruner {
|
||||
if cfg.Registry.URL == "" {
|
||||
return nil
|
||||
}
|
||||
token := os.Getenv(platform.RegistryPruneTokenEnv)
|
||||
if token == "" {
|
||||
fmt.Fprintf(stderr, "felis api: registry pruner disabled (%s unset) — images nothing uses are never deleted from the registry\n", platform.RegistryPruneTokenEnv)
|
||||
return nil
|
||||
}
|
||||
static := append([]string{os.Getenv("FELIS_IMAGE")}, buildConfig(cfg).ToolRefs()...)
|
||||
return ®istryprune.Pruner{
|
||||
Registry: ®istryprune.Client{Endpoint: "http://" + cfg.Registry.URL, Token: token},
|
||||
Host: cfg.Registry.URL,
|
||||
Refs: func(ctx context.Context) ([]string, error) {
|
||||
return inUseImageRefs(ctx, store, servers, static)
|
||||
},
|
||||
Log: slog.New(slog.NewTextHandler(stderr, nil)),
|
||||
}
|
||||
}
|
||||
|
||||
type imageRefStore interface {
|
||||
ListImages(ctx context.Context) ([]build.Image, error)
|
||||
ListUnfinishedBuilds(ctx context.Context) ([]build.Build, error)
|
||||
}
|
||||
|
||||
type serverLister interface {
|
||||
ListServers(ctx context.Context) ([]api.ServerInfo, error)
|
||||
PodImages(ctx context.Context) ([]string, error)
|
||||
}
|
||||
|
||||
// inUseImageRefs lists every image reference the platform still depends on: the
|
||||
// whitelist (disabled rows too, an admin may enable them again), every server's
|
||||
// spec, the images the game pods run, builds still running, and the images the
|
||||
// control plane and the build Jobs run. Any source failing fails the whole list,
|
||||
// so the pruner never decides on a partial view.
|
||||
//
|
||||
// The pods matter for the felis image: a running server keeps the one it started
|
||||
// with across platform upgrades (operator.PodTemplateAnnotation), which after a
|
||||
// few releases is no longer among the newest tags the pruner keeps anyway, and
|
||||
// the pod needs it again whenever it is recreated.
|
||||
func inUseImageRefs(ctx context.Context, store imageRefStore, servers serverLister, static []string) ([]string, error) {
|
||||
refs := append([]string(nil), static...)
|
||||
images, err := store.ListImages(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("image whitelist: %w", err)
|
||||
}
|
||||
for _, img := range images {
|
||||
refs = append(refs, img.ImageRef)
|
||||
}
|
||||
srvs, err := servers.ListServers(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("servers: %w", err)
|
||||
}
|
||||
for _, s := range srvs {
|
||||
refs = append(refs, s.Image)
|
||||
}
|
||||
podImages, err := servers.PodImages(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("game pods: %w", err)
|
||||
}
|
||||
refs = append(refs, podImages...)
|
||||
builds, err := store.ListUnfinishedBuilds(ctx)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("running builds: %w", err)
|
||||
}
|
||||
for _, b := range builds {
|
||||
refs = append(refs, b.ImageRef)
|
||||
}
|
||||
return refs, nil
|
||||
}
|
||||
|
||||
// mailLimit turns smtp.max_per_hour into the API's install-wide mail bucket:
|
||||
// the hourly cap as the refill rate, with a quarter of it (at least 5) allowed
|
||||
// at once so a burst of real sign-ins is not queued behind the average.
|
||||
func mailLimit(perHour int) api.RateLimit {
|
||||
if perHour <= 0 {
|
||||
perHour = config.DefaultMailPerHour
|
||||
}
|
||||
return api.RateLimit{Burst: max(perHour/4, 5), PerMinute: float64(perHour) / 60}
|
||||
}
|
||||
|
||||
// imagePinner resolves a new server's image against the platform registry
|
||||
// through its in-cluster Service, the address its refs already spell. An install
|
||||
// without a registry has no platform-built images to pin.
|
||||
func imagePinner(registry string) api.ImagePinner {
|
||||
if registry == "" {
|
||||
return nil
|
||||
}
|
||||
return imagepin.Resolver{Registry: registry}
|
||||
}
|
||||
@@ -1,10 +1,80 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/http"
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/api"
|
||||
"felis.lolicon.best/internal/build"
|
||||
"felis.lolicon.best/internal/config"
|
||||
)
|
||||
|
||||
// TestAuthSourcesFromConfig pins the one place the hasJoined identity anchor is decided:
|
||||
// Mojang is prepended in code, first, and is the only source whose UUIDs are trusted as-is.
|
||||
// The empty case matters on its own — both `felis api` and `felis nano` call this with a
|
||||
// config that has no [[auth_source]] at all, and that has to be a Mojang-only relay rather
|
||||
// than an empty list that rejects every login.
|
||||
func TestAuthSourcesFromConfig(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
configured []config.AuthSourceConfig
|
||||
}{
|
||||
{"no configured sources", nil},
|
||||
{"configured sources", []config.AuthSourceConfig{
|
||||
{Tag: "littleskin", Prefix: "LS", URL: "https://littleskin.example/hasJoined"},
|
||||
{Tag: "guild", Prefix: "GD", URL: "https://guild.example/hasJoined"},
|
||||
}},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
got := authSourcesFromConfig(tc.configured)
|
||||
if len(got) != len(tc.configured)+1 {
|
||||
t.Fatalf("got %d sources, want Mojang + %d configured", len(got), len(tc.configured))
|
||||
}
|
||||
if got[0].Tag != "mojang" || got[0].URL != mojangSessionServer || !got[0].Identity {
|
||||
t.Errorf("first source = %+v, want the Mojang identity anchor", got[0])
|
||||
}
|
||||
for i, c := range tc.configured {
|
||||
s := got[i+1]
|
||||
if s.Identity {
|
||||
t.Errorf("configured source %q is marked Identity; only Mojang may be", c.Tag)
|
||||
}
|
||||
if s.Tag != c.Tag || s.Prefix != c.Prefix || s.URL != c.URL {
|
||||
t.Errorf("source %d = %+v, want %+v in config order", i+1, s, c)
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestBuildConfig_ProjectsOverrides pins the [registry] overrides reaching the
|
||||
// build subsystem: unset fields must stay EMPTY (the build package's compiled-in
|
||||
// defaults apply there, not here), and set fields must pass through verbatim —
|
||||
// an air-gapped install points these at its imported mirrors.
|
||||
func TestBuildConfig_ProjectsOverrides(t *testing.T) {
|
||||
empty := buildConfig(&config.Config{})
|
||||
if empty.KanikoImage != "" || empty.TrivyImage != "" || empty.CPULimit != "" || empty.MemLimit != "" {
|
||||
t.Errorf("empty registry config must project empty overrides (defaults live in internal/build), got %+v", empty)
|
||||
}
|
||||
full := buildConfig(&config.Config{Registry: config.RegistryConfig{
|
||||
URL: "registry.felis.svc:5000",
|
||||
BuildNamespace: "felis-build",
|
||||
KanikoImage: "reg/kaniko:v1",
|
||||
TrivyImage: "reg/trivy:v1",
|
||||
BuildCPULimit: "1",
|
||||
BuildMemLimit: "2Gi",
|
||||
}})
|
||||
if full.KanikoImage != "reg/kaniko:v1" || full.TrivyImage != "reg/trivy:v1" ||
|
||||
full.CPULimit != "1" || full.MemLimit != "2Gi" {
|
||||
t.Errorf("registry overrides did not reach build.Config: %+v", full)
|
||||
}
|
||||
if full.Namespace != "felis-build" || full.RegistryURL != "registry.felis.svc:5000" {
|
||||
t.Errorf("namespace/registry url must keep projecting: %+v", full)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewAPIServerSetsHardenedTimeouts pins the gosec-G112 hardening on every
|
||||
// felis-api listener: the shared factory must bound the header and idle phases
|
||||
// (Slowloris + idle-connection exhaustion) while leaving WriteTimeout UNSET, because
|
||||
@@ -26,3 +96,62 @@ func TestNewAPIServerSetsHardenedTimeouts(t *testing.T) {
|
||||
t.Errorf("ReadTimeout = %v, want 0 (unset) so a slow SSE attach is not capped", srv.ReadTimeout)
|
||||
}
|
||||
}
|
||||
|
||||
type fakeRefStore struct {
|
||||
images []build.Image
|
||||
builds []build.Build
|
||||
err error
|
||||
}
|
||||
|
||||
func (f fakeRefStore) ListImages(context.Context) ([]build.Image, error) { return f.images, f.err }
|
||||
func (f fakeRefStore) ListUnfinishedBuilds(context.Context) ([]build.Build, error) {
|
||||
return f.builds, nil
|
||||
}
|
||||
|
||||
type fakeServers struct {
|
||||
list []api.ServerInfo
|
||||
pods []string
|
||||
podsErr error
|
||||
}
|
||||
|
||||
func (f fakeServers) ListServers(context.Context) ([]api.ServerInfo, error) { return f.list, nil }
|
||||
func (f fakeServers) PodImages(context.Context) ([]string, error) { return f.pods, f.podsErr }
|
||||
|
||||
// The registry pruner deletes whatever this list does not name, so every source of
|
||||
// a reference has to be in it, and a failing source must fail the list.
|
||||
func TestInUseImageRefsCoversEverySource(t *testing.T) {
|
||||
const reg = "registry.felis.svc:5000/"
|
||||
store := fakeRefStore{
|
||||
images: []build.Image{{ImageRef: reg + "modpacks/pack:*"}, {ImageRef: reg + "felis/paper:demo"}},
|
||||
builds: []build.Build{{ImageRef: reg + "user-uploads/sub-9:latest"}},
|
||||
}
|
||||
servers := fakeServers{
|
||||
list: []api.ServerInfo{{Name: "s1", Image: reg + "felis/paper:demo@sha256:" + fmt.Sprintf("%064d", 1)}},
|
||||
pods: []string{reg + "felis/felis:v1.0.0"},
|
||||
}
|
||||
got, err := inUseImageRefs(context.Background(), store, servers, []string{reg + "felis/felis:b60"})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
want := []string{
|
||||
reg + "felis/felis:b60",
|
||||
reg + "modpacks/pack:*", reg + "felis/paper:demo",
|
||||
reg + "felis/paper:demo@sha256:" + fmt.Sprintf("%064d", 1),
|
||||
reg + "felis/felis:v1.0.0",
|
||||
reg + "user-uploads/sub-9:latest",
|
||||
}
|
||||
if fmt.Sprint(got) != fmt.Sprint(want) {
|
||||
t.Fatalf("refs = %v\nwant %v", got, want)
|
||||
}
|
||||
|
||||
servers.podsErr = errors.New("apiserver down")
|
||||
if _, err := inUseImageRefs(context.Background(), store, servers, nil); err == nil {
|
||||
t.Fatal("a failing pod list produced a reference list")
|
||||
}
|
||||
servers.podsErr = nil
|
||||
|
||||
store.err = errors.New("db down")
|
||||
if _, err := inUseImageRefs(context.Background(), store, servers, nil); err == nil {
|
||||
t.Fatal("a failing whitelist read produced a reference list")
|
||||
}
|
||||
}
|
||||
+14
-1
@@ -220,7 +220,7 @@ func buildMinecraftServerFromApplyRequest(req applyRequest, namespace string) (*
|
||||
}
|
||||
memLim, ok := limits[corev1.ResourceMemory]
|
||||
if !ok || memLim.IsZero() {
|
||||
return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling (§22)")
|
||||
return nil, fmt.Errorf("internal error: refusing to create a server without a memory ceiling")
|
||||
}
|
||||
|
||||
// ---- storage ----
|
||||
@@ -243,6 +243,19 @@ func buildMinecraftServerFromApplyRequest(req applyRequest, namespace string) (*
|
||||
AutostartPolicy: policy,
|
||||
Storage: v1alpha1.StorageSpec{Size: storageQ.String()},
|
||||
Resources: corev1.ResourceRequirements{Limits: limits, Requests: requests},
|
||||
// The rest matches what felis-api's create writes (K8sCluster.CreateServer):
|
||||
// fall back to the login gate while stopped, RCON on (readiness, the
|
||||
// player count and the console all ride it; the operator mints the
|
||||
// password), and the default idle stop.
|
||||
FallbackServer: naming.SystemLoginServer,
|
||||
Rcon: v1alpha1.RconSpec{
|
||||
Enabled: true,
|
||||
SecretRef: v1alpha1.SecretKeyRef{
|
||||
Name: naming.RconSecretName(req.Name),
|
||||
Key: naming.RconSecretKey,
|
||||
},
|
||||
},
|
||||
Idle: v1alpha1.DefaultIdle(),
|
||||
},
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -6,6 +6,7 @@ import (
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/naming"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
"k8s.io/apimachinery/pkg/api/resource"
|
||||
)
|
||||
@@ -167,6 +168,18 @@ func TestBuildMinecraftServerFromApplyRequest_Valid(t *testing.T) {
|
||||
if ms.Spec.Storage.Size != "20Gi" {
|
||||
t.Errorf("Storage.Size = %q, want 20Gi", ms.Spec.Storage.Size)
|
||||
}
|
||||
// Same operational defaults as the API create path: without RCON the server
|
||||
// never reports players and the console answers 503; without spec.idle it
|
||||
// never stops on its own.
|
||||
if !ms.Spec.Rcon.Enabled || ms.Spec.Rcon.SecretRef.Name != naming.RconSecretName("test-server") {
|
||||
t.Errorf("Rcon = %+v, want enabled with the operator-minted secret", ms.Spec.Rcon)
|
||||
}
|
||||
if ms.Spec.Idle != v1alpha1.DefaultIdle() {
|
||||
t.Errorf("Idle = %+v, want the default %+v", ms.Spec.Idle, v1alpha1.DefaultIdle())
|
||||
}
|
||||
if ms.Spec.FallbackServer != naming.SystemLoginServer {
|
||||
t.Errorf("FallbackServer = %q, want the login gate", ms.Spec.FallbackServer)
|
||||
}
|
||||
mem, ok := ms.Spec.Resources.Limits[corev1.ResourceMemory]
|
||||
if !ok {
|
||||
t.Fatal("memory limit missing")
|
||||
|
||||
+38
-5
@@ -1,6 +1,7 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"crypto/rand"
|
||||
"encoding/hex"
|
||||
"flag"
|
||||
@@ -52,8 +53,8 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
|
||||
fmt.Fprintf(stderr, "felis backup: archive store %q is not implemented in this build (only tarLocal)\n", cfg.Archive.Store)
|
||||
return 1
|
||||
}
|
||||
// Reuse the reaper's retention derivation so an on-demand backup expires on the
|
||||
// same clock as an inactivity backup — one retention policy, not two.
|
||||
// The [archive] parse the reaper uses; an on-demand backup takes its
|
||||
// manual_retention and manual_keep.
|
||||
rcfg, err := reaperConfig(cfg)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: %v\n", err)
|
||||
@@ -72,6 +73,13 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
|
||||
|
||||
ctx := ctrl.SetupSignalHandler()
|
||||
|
||||
// The archive store shares the node's disk with every world and the
|
||||
// database: an owner's backup must not be what tips it into eviction.
|
||||
if err := backup.CheckRoom(cfg.Archive.LocalPath, *worldsRoot, backup.MinFreeAfter); err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
ref, size, err := archiver.Archive(ctx, *server, naming.WorldPVCName(*server))
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: archive: %v\n", err)
|
||||
@@ -92,9 +100,10 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
|
||||
BackupRef: string(ref),
|
||||
SizeBytes: size,
|
||||
Reason: "manual",
|
||||
ExpiresAt: time.Now().Add(rcfg.Retention),
|
||||
ExpiresAt: time.Now().Add(rcfg.ManualRetention),
|
||||
}
|
||||
if err := reaper.NewPGStore(drv.DB()).InsertBackup(ctx, rec); err != nil {
|
||||
st := reaper.NewPGStore(drv.DB())
|
||||
if err := st.InsertBackup(ctx, rec); err != nil {
|
||||
// The archive is written but unrecorded — an orphan the retention pass would
|
||||
// never expire. Delete it so a failed backup leaves no leaked bytes, mirroring
|
||||
// the reaper's archive-then-record atomicity.
|
||||
@@ -107,16 +116,40 @@ func cmdBackup(args []string, stdout, stderr io.Writer) int {
|
||||
}
|
||||
|
||||
fmt.Fprintf(stdout, "felis backup: server=%s archived %d bytes to %s (backup %s)\n", *server, size, ref, rec.ID)
|
||||
pruneManualBackups(ctx, st, archiver, *server, rcfg.ManualKeep, stdout, stderr)
|
||||
return 0
|
||||
}
|
||||
|
||||
// pruneManualBackups keeps server's newest keep on-demand backups and removes
|
||||
// the rest, oldest first, so repeated backups of one world cannot fill the
|
||||
// shared archive store. The new backup is already recorded; a removal that
|
||||
// fails is reported and retried after the next backup.
|
||||
func pruneManualBackups(ctx context.Context, st *reaper.PGStore, archiver backup.WorldArchiver, server string, keep int, stdout, stderr io.Writer) {
|
||||
excess, err := st.ExcessManualBackups(ctx, server, keep)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: list older backups of %s: %v\n", server, err)
|
||||
return
|
||||
}
|
||||
for _, b := range excess {
|
||||
if err := archiver.Delete(ctx, backup.ArchiveRef(b.BackupRef)); err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: remove older backup %s: %v\n", b.ID, err)
|
||||
continue
|
||||
}
|
||||
if err := st.MarkBackupDeleted(ctx, b.ID, time.Now()); err != nil {
|
||||
fmt.Fprintf(stderr, "felis backup: record the removal of %s: %v\n", b.ID, err)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis backup: removed older backup %s of %s (keeping the newest %d)\n", b.ID, server, keep)
|
||||
}
|
||||
}
|
||||
|
||||
// newBackupID mints a world_backups primary key, matching the reaper's "bk-"+hex
|
||||
// scheme so a manual and an inactivity backup are indistinguishable downstream.
|
||||
func newBackupID() string {
|
||||
var b [16]byte
|
||||
if _, err := rand.Read(b[:]); err != nil {
|
||||
// crypto/rand failure is fatal and unrecoverable; a time-based fallback would
|
||||
// be a weaker ID for no benefit. ponytail: panic is the honest failure here.
|
||||
// be a weaker ID for no benefit. A panic is the honest failure here.
|
||||
panic("felis backup: crypto/rand: " + err.Error())
|
||||
}
|
||||
return "bk-" + hex.EncodeToString(b[:])
|
||||
|
||||
+34
-10
@@ -51,7 +51,7 @@ func resolveInternalAPI(ctx context.Context, cl client.Client, controlNamespace
|
||||
}
|
||||
token = string(sec.Data[naming.ServiceTokenSecretKey])
|
||||
if token == "" {
|
||||
return "", "", fmt.Errorf("Secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey)
|
||||
return "", "", fmt.Errorf("secret %s has no %s key", naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey)
|
||||
}
|
||||
|
||||
return fmt.Sprintf("http://%s:%d", ip, platform.APIInternalPort), token, nil
|
||||
@@ -86,23 +86,31 @@ func requestBackup(ctx context.Context, hc *http.Client, baseURL, token, name, o
|
||||
|
||||
// backupErrorFromResponse turns a non-202 into a human message. The well-known codes get
|
||||
// an operator-facing explanation; anything else falls back to the API's
|
||||
// {"error":{message}} body, then the bare status code.
|
||||
// {"error":{code,message}} body, then the bare status code.
|
||||
func backupErrorFromResponse(resp *http.Response) error {
|
||||
switch resp.StatusCode {
|
||||
case http.StatusConflict: // not_stopped
|
||||
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
|
||||
case http.StatusServiceUnavailable: // backup_unavailable
|
||||
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
|
||||
case http.StatusNotFound:
|
||||
return fmt.Errorf("no such server")
|
||||
}
|
||||
var e struct {
|
||||
Error struct {
|
||||
Code string `json:"code"`
|
||||
Message string `json:"message"`
|
||||
} `json:"error"`
|
||||
}
|
||||
raw, _ := io.ReadAll(io.LimitReader(resp.Body, 1<<16))
|
||||
_ = json.Unmarshal(raw, &e)
|
||||
|
||||
switch resp.StatusCode {
|
||||
case http.StatusConflict:
|
||||
// Two refusals share 409: the stopped gate and the missing-world-volume
|
||||
// gate. The body's code distinguishes them; a code-less body reads as the
|
||||
// stopped gate (the only 409 before the volume gate existed), and any other
|
||||
// coded 409 falls through to the API's own operator text.
|
||||
if e.Error.Code == "" || e.Error.Code == "not_stopped" {
|
||||
return fmt.Errorf("the server must be stopped before its world can be backed up — halt it first")
|
||||
}
|
||||
case http.StatusServiceUnavailable: // backup_unavailable
|
||||
return fmt.Errorf("the backup subsystem is not configured on felis-api (FELIS_IMAGE / FELIS_BACKUP_PVC unset)")
|
||||
case http.StatusNotFound:
|
||||
return fmt.Errorf("no such server")
|
||||
}
|
||||
if e.Error.Message != "" {
|
||||
return fmt.Errorf("felis-api: %s", e.Error.Message)
|
||||
}
|
||||
@@ -120,3 +128,19 @@ func performBackupNow(ctx context.Context, cl client.Client, controlNamespace, n
|
||||
hc := &http.Client{Timeout: 10 * time.Second}
|
||||
return requestBackup(ctx, hc, baseURL, token, name, osUser)
|
||||
}
|
||||
|
||||
// backupPickable narrows the backup picker to servers the backup API can accept.
|
||||
// System servers (login/lobby) are excluded: they have no row in the servers
|
||||
// table and carry reserved names, so every attempt dies in name validation —
|
||||
// offering them would be a dead pick. The halt picker keeps them on purpose
|
||||
// (break-glass retains full power over system servers); only the API-backed
|
||||
// backup op cannot reach them.
|
||||
func backupPickable(servers []haltableServer) []haltableServer {
|
||||
out := make([]haltableServer, 0, len(servers))
|
||||
for _, s := range servers {
|
||||
if !s.system {
|
||||
out = append(out, s)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -114,16 +114,23 @@ func TestRequestBackup(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
code int
|
||||
body string // optional JSON error body
|
||||
expect string
|
||||
}{
|
||||
{"409 not_stopped", http.StatusConflict, "must be stopped"},
|
||||
{"503 backup_unavailable", http.StatusServiceUnavailable, "not configured"},
|
||||
{"404 not found", http.StatusNotFound, "no such server"},
|
||||
{"409 not_stopped", http.StatusConflict, "", "must be stopped"},
|
||||
{"409 no_world_volume surfaces the API's own text", http.StatusConflict,
|
||||
`{"error":{"code":"no_world_volume","message":"this server has no world volume yet — start it once to create it, then retry"}}`,
|
||||
"no world volume yet"},
|
||||
{"503 backup_unavailable", http.StatusServiceUnavailable, "", "not configured"},
|
||||
{"404 not found", http.StatusNotFound, "", "no such server"},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
w.WriteHeader(tc.code)
|
||||
if tc.body != "" {
|
||||
_, _ = io.WriteString(w, tc.body)
|
||||
}
|
||||
}))
|
||||
defer srv.Close()
|
||||
_, err := requestBackup(context.Background(), hc, srv.URL, "tok", "survival", "alice")
|
||||
@@ -142,3 +149,25 @@ func TestRequestBackup(t *testing.T) {
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// The Sync picker must not offer system servers: the backup API validates names
|
||||
// and resolves a servers-table row, so a lobby/login pick can only die in
|
||||
// validation — a dead choice in an emergency console.
|
||||
func TestBackupPickable(t *testing.T) {
|
||||
got := backupPickable([]haltableServer{
|
||||
{name: "lobby", phase: "Running", system: true},
|
||||
{name: "login", phase: "Running", system: true},
|
||||
{name: "test-one", phase: "Stopped"},
|
||||
})
|
||||
if len(got) != 1 || got[0].name != "test-one" || got[0].system {
|
||||
t.Fatalf("backupPickable = %+v, want only the user server", got)
|
||||
}
|
||||
|
||||
// Survivors keep their input order (the picker's cursor math depends on it).
|
||||
got = backupPickable([]haltableServer{
|
||||
{name: "alpha"}, {name: "login", system: true}, {name: "beta"},
|
||||
})
|
||||
if len(got) != 2 || got[0].name != "alpha" || got[1].name != "beta" {
|
||||
t.Fatalf("backupPickable order = %+v, want [alpha beta]", got)
|
||||
}
|
||||
}
|
||||
+47
-18
@@ -76,12 +76,19 @@ type ownerStore interface {
|
||||
AdminExists(ctx context.Context) (bool, error)
|
||||
// UserByUsername loads a staff login projection.
|
||||
UserByUsername(ctx context.Context, username string) (*api.StaffUser, error)
|
||||
// OwnerUsername names the single active Owner seat, or "" when none exists.
|
||||
// provisionOwner refuses to re-target anything but this username: with the
|
||||
// seat occupied, a fresh name would mint a second owner row (UpsertOwner's
|
||||
// insert arm) while the existing — possibly compromised — seat stays live,
|
||||
// and no supported path can delete an owner row afterwards.
|
||||
OwnerUsername(ctx context.Context) (string, error)
|
||||
UpsertOwner(ctx context.Context, id, username, email string) error
|
||||
// InsertOperator mints a NEW Operator staff account. Unlike UpsertOwner it is
|
||||
// insert-only: a username already taken is a conflict (api.ErrConflict), never a
|
||||
// silent reset, so adding an Operator can never clobber the Owner or an existing
|
||||
// Operator. The row is role=admin, identical in shape to the Owner — Felis has no
|
||||
// separate operator DB role (migration 0003: staff = role=admin).
|
||||
// Operator. The row is role=admin — an Operator is staff BELOW the single
|
||||
// role=owner identity (migration 0011 adds that role); the two are the only
|
||||
// staff roles.
|
||||
InsertOperator(ctx context.Context, id, username, email string) error
|
||||
// CompleteOwnerSetup atomically consumes the in-game link code, creates or
|
||||
// promotes the bound Owner, enables local auth, and stores the one-time setup
|
||||
@@ -266,20 +273,31 @@ func authenticateAdmin(ctx context.Context, s ownerStore, username string) (matc
|
||||
if err != nil {
|
||||
return "", false, err
|
||||
}
|
||||
if u.Role != "admin" {
|
||||
// Staff means admin OR owner: recovery attribution must accept the Owner (the
|
||||
// primary break-glass identity), not just plain admins.
|
||||
if u.Role != "admin" && u.Role != "owner" {
|
||||
return "", false, nil
|
||||
}
|
||||
return u.Username, true, nil
|
||||
}
|
||||
|
||||
// provisionOwner mints or resets the single Owner account direct-to-Postgres,
|
||||
// passwordless. The account is role=admin with no password — the Owner completes
|
||||
// passwordless. The account is role=owner with no password — the Owner completes
|
||||
// passwordless login setup via the web setup-token flow after `felis setup`.
|
||||
// With a seat already occupied the reset must name that seat (ownerSeatTakenError
|
||||
// otherwise): the upsert's insert arm would silently mint a SECOND owner, and
|
||||
// every owner row is undeletable through the panel, so the tier could never
|
||||
// converge back to one.
|
||||
func provisionOwner(ctx context.Context, s ownerStore, username, email string) error {
|
||||
username = strings.TrimSpace(username)
|
||||
if username == "" {
|
||||
return errors.New("owner username is required")
|
||||
}
|
||||
if seat, err := s.OwnerUsername(ctx); err != nil {
|
||||
return fmt.Errorf("check the owner seat: %w", err)
|
||||
} else if seat != "" && seat != username {
|
||||
return &ownerSeatTakenError{seat: seat}
|
||||
}
|
||||
id := newOwnerID()
|
||||
if id == "" {
|
||||
return errors.New("generate owner id: entropy source failed")
|
||||
@@ -290,13 +308,26 @@ func provisionOwner(ctx context.Context, s ownerStore, username, email string) e
|
||||
return nil
|
||||
}
|
||||
|
||||
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. Like the
|
||||
// Owner it is role=admin and passwordless — Felis has no separate operator DB role,
|
||||
// so an Operator is simply an additional staff admin (migration 0003). UNLIKE
|
||||
// provisionOwner, which upserts the single Owner and resets it on a username
|
||||
// conflict, this is insert-only: a username already taken returns api.ErrConflict
|
||||
// rather than overwriting a live account, so adding an Operator can never silently
|
||||
// clobber the Owner's or another Operator's account.
|
||||
// ownerSeatTakenError refuses an Owner reset that names anything but the
|
||||
// occupied seat, naming it so the operator can retype. Is reports
|
||||
// api.ErrConflict so the TUI's recoverable-error branch (shared with the
|
||||
// operator path's taken-name clash) routes back to the form instead of ending
|
||||
// the console.
|
||||
type ownerSeatTakenError struct{ seat string }
|
||||
|
||||
func (e *ownerSeatTakenError) Error() string {
|
||||
return fmt.Sprintf("an Owner already exists as %q — enter that username to reset the Owner", e.seat)
|
||||
}
|
||||
|
||||
func (e *ownerSeatTakenError) Is(target error) bool { return target == api.ErrConflict }
|
||||
|
||||
// provisionOperator mints a NEW Operator staff account direct-to-Postgres. It is
|
||||
// role=admin and passwordless — an additional staff admin below the single
|
||||
// role=owner identity (migrations 0003 + 0011). UNLIKE provisionOwner, which
|
||||
// upserts the single Owner and resets it on a username conflict, this is
|
||||
// insert-only: a username already taken returns api.ErrConflict rather than
|
||||
// overwriting a live account, so adding an Operator can never silently clobber
|
||||
// the Owner's or another Operator's account.
|
||||
func provisionOperator(ctx context.Context, s ownerStore, username, email string) error {
|
||||
username = strings.TrimSpace(username)
|
||||
if username == "" {
|
||||
@@ -387,7 +418,7 @@ func newSetupToken() (raw, hash string, err error) {
|
||||
|
||||
// performSetupMCBind is the `felis setup` Owner-establishment path: the operator
|
||||
// binds their Minecraft account via a one-time link code the login gate handed
|
||||
// them in-game, the bound user is promoted to role='admin' (passwordless Owner),
|
||||
// them in-game, the bound user is promoted to role='owner' (passwordless Owner),
|
||||
// local auth is enabled, and a one-time setup URL is minted for the first web
|
||||
// login where the Owner verifies email / enrolls a passkey. adminHostname is the
|
||||
// operator-console host the URL points at (op.console.<root>): the Owner is staff,
|
||||
@@ -571,12 +602,10 @@ type breakGlassResult struct {
|
||||
backupStatus string
|
||||
|
||||
// Cloudflare-specific edge detail (set only when connectMethod is Cloudflare)
|
||||
edgeConfigured bool
|
||||
edgeAud string
|
||||
edgeRoutedHosts []string
|
||||
edgeConfigPath string
|
||||
edgePanelHostname string
|
||||
edgeAdminHostname string
|
||||
edgeConfigured bool
|
||||
edgeAud string
|
||||
edgeRoutedHosts []string
|
||||
edgeConfigPath string
|
||||
}
|
||||
|
||||
type consoleMode string
|
||||
|
||||
@@ -19,14 +19,15 @@ import (
|
||||
// terminal. The design is passwordless: accounts carry no credential, and the
|
||||
// Owner completes first-login through the setup-token web flow.
|
||||
type fakeOwnerStore struct {
|
||||
upserts []upsertCall
|
||||
inserts []upsertCall
|
||||
settings map[string][]byte
|
||||
audits []api.AuditEntry
|
||||
tokens []setupTokenCall
|
||||
redeems []redeemCall
|
||||
users map[string]*api.StaffUser // keyed by username
|
||||
admins bool // AdminExists answer
|
||||
upserts []upsertCall
|
||||
inserts []upsertCall
|
||||
settings map[string][]byte
|
||||
audits []api.AuditEntry
|
||||
tokens []setupTokenCall
|
||||
redeems []redeemCall
|
||||
users map[string]*api.StaffUser // keyed by username
|
||||
admins bool // AdminExists answer
|
||||
ownerSeat string // OwnerUsername answer: the occupied seat, "" when none
|
||||
|
||||
// CompleteOwnerSetup's success result. redeemUserID defaults to the fresh id
|
||||
// the caller passes (the unlinked-UUID case) when left empty.
|
||||
@@ -40,6 +41,7 @@ type fakeOwnerStore struct {
|
||||
auditErr error
|
||||
userErr error // non-not-found error from UserByUsername
|
||||
adminErr error
|
||||
seatErr error
|
||||
redeemErr error
|
||||
createTokenErr error
|
||||
}
|
||||
@@ -80,6 +82,15 @@ func (f *fakeOwnerStore) UserByUsername(_ context.Context, username string) (*ap
|
||||
return nil, api.ErrNotFound
|
||||
}
|
||||
|
||||
// OwnerUsername reports the single active Owner seat. Tests set ownerSeat; the
|
||||
// zero value models a fresh install where bootstrap is free to mint.
|
||||
func (f *fakeOwnerStore) OwnerUsername(_ context.Context) (string, error) {
|
||||
if f.seatErr != nil {
|
||||
return "", f.seatErr
|
||||
}
|
||||
return f.ownerSeat, nil
|
||||
}
|
||||
|
||||
func (f *fakeOwnerStore) UpsertOwner(_ context.Context, id, username, email string) error {
|
||||
if f.upsertErr != nil {
|
||||
return f.upsertErr
|
||||
@@ -201,6 +212,34 @@ func TestProvisionOwner(t *testing.T) {
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("an occupied seat refuses any other username", func(t *testing.T) {
|
||||
// The seat is the single owner row: upserting a fresh name would take the
|
||||
// insert arm and mint a SECOND owner, while the existing seat — possibly the
|
||||
// compromised account this reset was meant to replace — stays live, and no
|
||||
// supported path can delete an owner row.
|
||||
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
|
||||
err := provisionOwner(ctx, f, "someone-else", "")
|
||||
if !errors.Is(err, api.ErrConflict) {
|
||||
t.Fatalf("error = %v, want it to wrap api.ErrConflict so the TUI routes back to the form", err)
|
||||
}
|
||||
if !strings.Contains(err.Error(), `"seat-holder"`) {
|
||||
t.Errorf("error = %q, want it to name the occupied seat", err)
|
||||
}
|
||||
if len(f.upserts) != 0 {
|
||||
t.Errorf("want no write against an occupied seat, got %d", len(f.upserts))
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("the occupied seat's own username still resets", func(t *testing.T) {
|
||||
f := &fakeOwnerStore{ownerSeat: "seat-holder"}
|
||||
if err := provisionOwner(ctx, f, "seat-holder", "[email protected]"); err != nil {
|
||||
t.Fatalf("provisionOwner(reset): %v", err)
|
||||
}
|
||||
if len(f.upserts) != 1 || f.upserts[0].username != "seat-holder" || f.upserts[0].email != "[email protected]" {
|
||||
t.Fatalf("want 1 reset upsert for the seat, got %+v", f.upserts)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("propagates a store error", func(t *testing.T) {
|
||||
f := &fakeOwnerStore{upsertErr: errors.New("boom")}
|
||||
if err := provisionOwner(ctx, f, "owner", ""); err == nil {
|
||||
@@ -264,6 +303,16 @@ func TestAuthenticateAdmin(t *testing.T) {
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("the owner role attributes like an admin", func(t *testing.T) {
|
||||
owner := mkAdmin("root")
|
||||
owner.Role = "owner" // the platform owner is staff too (migration 0011)
|
||||
f := &fakeOwnerStore{users: map[string]*api.StaffUser{"root": owner}}
|
||||
matched, ok, err := authenticateAdmin(ctx, f, "root")
|
||||
if err != nil || !ok || matched != "root" {
|
||||
t.Fatalf("authenticateAdmin(owner) = (%q, %v, %v), want (root, true, nil)", matched, ok, err)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("an unknown user is a non-match, not an error", func(t *testing.T) {
|
||||
f := &fakeOwnerStore{}
|
||||
_, ok, err := authenticateAdmin(ctx, f, "nobody")
|
||||
|
||||
@@ -0,0 +1,105 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"strings"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
)
|
||||
|
||||
// cmdConverge is the explicit convergence pass over already-installed system
|
||||
// servers (#1), plus the idle-stop default for user servers that predate it. Provisioning is create-if-absent, so a field the desired spec
|
||||
// gained after an install (spec.rcon, spec.startup.healthHTTPPort, a derived env
|
||||
// key) never reaches the existing CR — and nothing says so. This command fills
|
||||
// exactly those zero-value fields; see convergeSystemServers for the full contract
|
||||
// and why it is a separate, operator-timed step rather than part of setup.
|
||||
//
|
||||
// It reads the same host config as setup (the control plane's felis.toml) and
|
||||
// talks to the cluster with the local kubeconfig, so it must run as root on the
|
||||
// control-plane host.
|
||||
func cmdConverge(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("converge", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
cfgPath := fs.String("config", defaultSetupConfigPath, "path to felis.toml")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return 0
|
||||
}
|
||||
return 2
|
||||
}
|
||||
if os.Geteuid() != 0 {
|
||||
fmt.Fprintln(stderr, "felis converge: refused — converging needs the cluster credentials, so it must run as root (try: sudo felis converge)")
|
||||
return 1
|
||||
}
|
||||
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis converge: %v\n", err)
|
||||
fmt.Fprintln(stderr, "If this host was never installed, run `sudo felis setup` first.")
|
||||
return 1
|
||||
}
|
||||
cl, err := buildSystemServerClient()
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis converge: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
controlNS := platform.DefaultControlNamespace
|
||||
outcomes := convergeSystemServers(context.Background(), cl, cfg.K8s.Namespace,
|
||||
cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage,
|
||||
platform.InternalAPIBaseURL(controlNS), cfg.Server.RootDomain,
|
||||
defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
|
||||
|
||||
outcomes = append(outcomes, convergeUserServerIdle(context.Background(), cl, cfg.K8s.Namespace)...)
|
||||
|
||||
fmt.Fprintln(stdout, "felis converge: filling fields an installed server predates (operator-set values are never overwritten):")
|
||||
exit := 0
|
||||
for _, o := range outcomes {
|
||||
switch {
|
||||
case o.err != nil:
|
||||
fmt.Fprintf(stdout, " - %s: ERROR %v\n", o.name, o.err)
|
||||
exit = 1
|
||||
case len(o.changes) > 0:
|
||||
fmt.Fprintf(stdout, " - %s: updated (%s)\n", o.name, strings.Join(o.changes, ", "))
|
||||
default:
|
||||
fmt.Fprintf(stdout, " - %s: %s\n", o.name, o.skipped)
|
||||
}
|
||||
}
|
||||
return exit
|
||||
}
|
||||
|
||||
// convergeUserServerIdle gives every user server that predates the idle default
|
||||
// (spec.idle entirely unset) the default idle stop. A server whose idle stop was
|
||||
// turned off keeps a duration on its spec, so it is not "unset" and is left
|
||||
// alone; system servers never idle out and are skipped. Servers that already
|
||||
// carry a value produce no line, so a converged fleet prints nothing here.
|
||||
func convergeUserServerIdle(ctx context.Context, cl client.Client, namespace string) []systemServerOutcome {
|
||||
var list v1alpha1.MinecraftServerList
|
||||
if err := cl.List(ctx, &list, client.InNamespace(namespace)); err != nil {
|
||||
return []systemServerOutcome{{name: "user servers", err: fmt.Errorf("list servers: %w", err)}}
|
||||
}
|
||||
var out []systemServerOutcome
|
||||
for i := range list.Items {
|
||||
ms := &list.Items[i]
|
||||
if ms.Labels[v1alpha1.LabelSystemRole] != "" || ms.Spec.Idle != (v1alpha1.IdleSpec{}) {
|
||||
continue
|
||||
}
|
||||
patch := client.MergeFrom(ms.DeepCopy())
|
||||
ms.Spec.Idle = v1alpha1.DefaultIdle()
|
||||
if err := cl.Patch(ctx, ms, patch); err != nil {
|
||||
out = append(out, systemServerOutcome{name: ms.Name, err: fmt.Errorf("converge %s: %w", ms.Name, err)})
|
||||
continue
|
||||
}
|
||||
out = append(out, systemServerOutcome{name: ms.Name, available: true, updated: true,
|
||||
changes: []string{fmt.Sprintf("spec.idle (stop after %ds empty)", v1alpha1.DefaultEmptySecondsBeforeStop)}})
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,231 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/naming"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client/fake"
|
||||
)
|
||||
|
||||
// converge is the explicit pass over an installed system server whose CR predates
|
||||
// a field the desired spec has since gained (#1). It must fill exactly the
|
||||
// zero-valued whitelist fields and the derived env, and must not touch anything a
|
||||
// non-zero value already occupies — that is the operator's.
|
||||
func TestConvergeSystemServersFillsPredatedFields(t *testing.T) {
|
||||
scheme := newSystemServerScheme(t)
|
||||
ctx := context.Background()
|
||||
|
||||
// An old install: the lobby CR was created before the desired spec began
|
||||
// rendering spec.rcon, and the login CR before the HTTP readiness gate existed.
|
||||
// One derived env key is absent entirely (as if it were added later), and one
|
||||
// hand-added env var plus a non-whitelisted spec field must survive.
|
||||
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
|
||||
if err != nil {
|
||||
t.Fatalf("build lobby: %v", err)
|
||||
}
|
||||
lobby.Spec.Rcon = v1alpha1.RconSpec{}
|
||||
lobby.Spec.JavaMemory = "999Mi"
|
||||
|
||||
login, err := loginSystemServer("reg/limbo:1", "minecraft",
|
||||
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
|
||||
if err != nil {
|
||||
t.Fatalf("build login: %v", err)
|
||||
}
|
||||
login.Spec.Startup.HealthHTTPPort = 0
|
||||
kept := login.Spec.Env
|
||||
login.Spec.Env = nil
|
||||
for _, e := range kept {
|
||||
if e.Name != envPanelHostname {
|
||||
login.Spec.Env = append(login.Spec.Env, e)
|
||||
}
|
||||
}
|
||||
login.Spec.Env = append(login.Spec.Env, v1alpha1.EnvVar{Name: "OPERATOR_TUNING", Value: "keep-me"})
|
||||
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
|
||||
outcomes := convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
|
||||
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
|
||||
|
||||
byName := map[string]systemServerOutcome{}
|
||||
for _, o := range outcomes {
|
||||
if o.err != nil {
|
||||
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
|
||||
}
|
||||
byName[o.name] = o
|
||||
}
|
||||
lobbyOut := byName[naming.SystemLobbyServer]
|
||||
if len(lobbyOut.changes) != 1 || lobbyOut.changes[0] != "spec.rcon" {
|
||||
t.Errorf("lobby changes = %v, want [spec.rcon] (only the zero-valued field)", lobbyOut.changes)
|
||||
}
|
||||
loginOut := byName[naming.SystemLoginServer]
|
||||
if !slices.Contains(loginOut.changes, "spec.startup.healthHTTPPort") || !slices.Contains(loginOut.changes, "env "+envPanelHostname) {
|
||||
t.Errorf("login changes = %v, want the health port plus the missing derived env key", loginOut.changes)
|
||||
}
|
||||
|
||||
var gotLobby v1alpha1.MinecraftServer
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &gotLobby); err != nil {
|
||||
t.Fatalf("get lobby: %v", err)
|
||||
}
|
||||
if !gotLobby.Spec.Rcon.Enabled ||
|
||||
gotLobby.Spec.Rcon.SecretRef.Name != naming.RconSecretName(naming.SystemLobbyServer) ||
|
||||
gotLobby.Spec.Rcon.SecretRef.Key != naming.RconSecretKey {
|
||||
t.Errorf("lobby rcon = %+v, want the desired block with the %s secret",
|
||||
gotLobby.Spec.Rcon, naming.RconSecretName(naming.SystemLobbyServer))
|
||||
}
|
||||
if gotLobby.Spec.JavaMemory != "999Mi" {
|
||||
t.Errorf("lobby javaMemory = %q, want 999Mi — converge fills new fields, it does not rewrite the spec", gotLobby.Spec.JavaMemory)
|
||||
}
|
||||
|
||||
var gotLogin v1alpha1.MinecraftServer
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLoginServer}, &gotLogin); err != nil {
|
||||
t.Fatalf("get login: %v", err)
|
||||
}
|
||||
if gotLogin.Spec.Startup.HealthHTTPPort != felisLimboHealthPort {
|
||||
t.Errorf("login healthHTTPPort = %d, want %d", gotLogin.Spec.Startup.HealthHTTPPort, felisLimboHealthPort)
|
||||
}
|
||||
env := map[string]string{}
|
||||
for _, e := range gotLogin.Spec.Env {
|
||||
env[e.Name] = e.Value
|
||||
}
|
||||
if env[envPanelHostname] != "console.mc.example.net" {
|
||||
t.Errorf("%s was not added back: %q", envPanelHostname, env[envPanelHostname])
|
||||
}
|
||||
if env["OPERATOR_TUNING"] != "keep-me" {
|
||||
t.Error("a hand-added env var was dropped; converge only touches config-derived names")
|
||||
}
|
||||
}
|
||||
|
||||
// A field already holding a non-zero value belongs to the operator: converge must
|
||||
// report "already converged" and write nothing.
|
||||
func TestConvergeSystemServersLeavesNonZeroFieldsAlone(t *testing.T) {
|
||||
scheme := newSystemServerScheme(t)
|
||||
ctx := context.Background()
|
||||
|
||||
lobby, err := lobbySystemServer("reg/lobby:1", "minecraft")
|
||||
if err != nil {
|
||||
t.Fatalf("build lobby: %v", err)
|
||||
}
|
||||
lobby.Spec.Rcon = v1alpha1.RconSpec{
|
||||
Enabled: true,
|
||||
SecretRef: v1alpha1.SecretKeyRef{Name: "operator-rotated", Key: "password"},
|
||||
}
|
||||
login, err := loginSystemServer("reg/limbo:1", "minecraft",
|
||||
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
|
||||
if err != nil {
|
||||
t.Fatalf("build login: %v", err)
|
||||
}
|
||||
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(lobby, login).Build()
|
||||
for _, o := range convergeSystemServers(ctx, cl, "minecraft", "reg/limbo:1", "reg/lobby:1",
|
||||
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net") {
|
||||
if o.err != nil {
|
||||
t.Fatalf("%s: unexpected error: %v", o.name, o.err)
|
||||
}
|
||||
if len(o.changes) != 0 || o.skipped != "already converged" {
|
||||
t.Errorf("%s outcome = %+v, want already converged with no writes", o.name, o)
|
||||
}
|
||||
}
|
||||
var got v1alpha1.MinecraftServer
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: naming.SystemLobbyServer}, &got); err != nil {
|
||||
t.Fatalf("get lobby: %v", err)
|
||||
}
|
||||
if got.Spec.Rcon.SecretRef.Name != "operator-rotated" {
|
||||
t.Errorf("lobby rcon secretRef = %q — converge overwrote a field the operator had already set",
|
||||
got.Spec.Rcon.SecretRef.Name)
|
||||
}
|
||||
}
|
||||
|
||||
// Guards: an absent CR is reported (creation is setup's job), a foreign CR is
|
||||
// refused rather than adopted, and an unset image skips like the provisioner does.
|
||||
func TestConvergeSystemServersGuards(t *testing.T) {
|
||||
scheme := newSystemServerScheme(t)
|
||||
ctx := context.Background()
|
||||
run := func(cl client.Client, loginImage, lobbyImage string) []systemServerOutcome {
|
||||
return convergeSystemServers(ctx, cl, "minecraft", loginImage, lobbyImage,
|
||||
"http://felis-api.felis.svc.cluster.local:8081", "mc.example.net", "console.mc.example.net")
|
||||
}
|
||||
|
||||
t.Run("absent CRs are reported, not created", func(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
|
||||
for _, o := range run(cl, "reg/limbo:1", "reg/lobby:1") {
|
||||
if o.err != nil {
|
||||
t.Fatalf("%s: %v", o.name, o.err)
|
||||
}
|
||||
if o.created || !strings.Contains(o.skipped, "not present") {
|
||||
t.Errorf("%s outcome = %+v, want a not-present skip", o.name, o)
|
||||
}
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("foreign CR is refused", func(t *testing.T) {
|
||||
foreign := &v1alpha1.MinecraftServer{}
|
||||
foreign.Name = naming.SystemLoginServer
|
||||
foreign.Namespace = "minecraft"
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(foreign).Build()
|
||||
out := run(cl, "reg/limbo:1", "")
|
||||
if len(out) != 2 {
|
||||
t.Fatalf("outcomes = %d, want 2", len(out))
|
||||
}
|
||||
if out[0].err == nil || !strings.Contains(out[0].err.Error(), "not marked") {
|
||||
t.Fatalf("login error = %v, want an unmarked-name refusal", out[0].err)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("unset image skips", func(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).Build()
|
||||
out := run(cl, "", "reg/lobby:1")
|
||||
if out[0].skipped != "image not configured" {
|
||||
t.Errorf("login skipped = %q, want %q", out[0].skipped, "image not configured")
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// TestConvergeUserServerIdle fills the idle default only where spec.idle was
|
||||
// never set: a server whose idle stop was turned off (duration kept), one with
|
||||
// its own duration, and a system server all stay as they are.
|
||||
func TestConvergeUserServerIdle(t *testing.T) {
|
||||
scheme := newSystemServerScheme(t)
|
||||
ctx := context.Background()
|
||||
mk := func(name string, idle v1alpha1.IdleSpec, role string) *v1alpha1.MinecraftServer {
|
||||
ms := &v1alpha1.MinecraftServer{}
|
||||
ms.Name, ms.Namespace = name, "minecraft"
|
||||
ms.Spec.Idle = idle
|
||||
if role != "" {
|
||||
ms.Labels = map[string]string{v1alpha1.LabelSystemRole: role}
|
||||
}
|
||||
return ms
|
||||
}
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
|
||||
mk("legacy", v1alpha1.IdleSpec{}, ""),
|
||||
mk("off", v1alpha1.IdleSpec{EmptySecondsBeforeStop: 600}, ""),
|
||||
mk("custom", v1alpha1.IdleSpec{AutoStopEnabled: true, EmptySecondsBeforeStop: 1800}, ""),
|
||||
mk(naming.SystemLobbyServer, v1alpha1.IdleSpec{}, naming.SystemLobbyServer),
|
||||
).Build()
|
||||
|
||||
outcomes := convergeUserServerIdle(ctx, cl, "minecraft")
|
||||
if len(outcomes) != 1 || outcomes[0].name != "legacy" || outcomes[0].err != nil {
|
||||
t.Fatalf("outcomes = %+v, want exactly one fill for legacy", outcomes)
|
||||
}
|
||||
want := map[string]v1alpha1.IdleSpec{
|
||||
"legacy": v1alpha1.DefaultIdle(),
|
||||
"off": {EmptySecondsBeforeStop: 600},
|
||||
"custom": {AutoStopEnabled: true, EmptySecondsBeforeStop: 1800},
|
||||
naming.SystemLobbyServer: {},
|
||||
}
|
||||
for name, idle := range want {
|
||||
var ms v1alpha1.MinecraftServer
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: name}, &ms); err != nil {
|
||||
t.Fatalf("get %s: %v", name, err)
|
||||
}
|
||||
if ms.Spec.Idle != idle {
|
||||
t.Errorf("%s idle = %+v, want %+v", name, ms.Spec.Idle, idle)
|
||||
}
|
||||
}
|
||||
if again := convergeUserServerIdle(ctx, cl, "minecraft"); len(again) != 0 {
|
||||
t.Fatalf("second pass = %+v, want nothing to do", again)
|
||||
}
|
||||
}
|
||||
+334
@@ -0,0 +1,334 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/dbbackup"
|
||||
)
|
||||
|
||||
const dbUsage = `usage:
|
||||
felis db backup [-config path] [-dir dir] [-label daily|manual|...] [-keep n] [-state-dir dir]
|
||||
[-no-servers] [-metrics-file path]
|
||||
felis db restore [-config path] [-dir dir] [-yes] [-force] [-no-safety-backup] <bundle>
|
||||
felis db verify [-dir dir] <bundle>
|
||||
felis db list [-dir dir]
|
||||
felis db check [-dir dir] [-max-age 26h]
|
||||
`
|
||||
|
||||
// defaultKeep is how many bundles of a label a backup leaves behind. Manual
|
||||
// bundles are the operator's own and are never pruned.
|
||||
var defaultKeep = map[string]int{
|
||||
dbbackup.LabelDaily: 14,
|
||||
dbbackup.LabelPreMigrate: 10,
|
||||
dbbackup.LabelPreRestore: 5,
|
||||
}
|
||||
|
||||
// cmdDB implements `felis db`: logical backups of the control-plane database
|
||||
// together with the host state a rebuild needs (internal/dbbackup). The verb
|
||||
// comes first for the same reason as `felis migrate up`.
|
||||
func cmdDB(args []string, stdout, stderr io.Writer) int {
|
||||
if len(args) == 0 {
|
||||
fmt.Fprint(stderr, dbUsage)
|
||||
return 2
|
||||
}
|
||||
verb, rest := args[0], args[1:]
|
||||
fs := flag.NewFlagSet("db "+verb, flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
fs.Usage = func() { fmt.Fprint(stderr, dbUsage) }
|
||||
dir := fs.String("dir", dbbackup.DefaultDir, "bundle directory")
|
||||
switch verb {
|
||||
case "backup":
|
||||
return dbBackup(fs, dir, rest, stdout, stderr)
|
||||
case "restore":
|
||||
return dbRestore(fs, dir, rest, stdout, stderr)
|
||||
case "verify":
|
||||
return dbVerify(fs, dir, rest, stdout, stderr)
|
||||
case "list":
|
||||
return dbList(fs, dir, rest, stdout, stderr)
|
||||
case "check":
|
||||
return dbCheck(fs, dir, rest, stdout, stderr)
|
||||
case "-h", "--help", "help":
|
||||
fmt.Fprint(stdout, dbUsage)
|
||||
return 0
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis db: unknown verb %q\n%s", verb, dbUsage)
|
||||
return 2
|
||||
}
|
||||
|
||||
// parseWithArg parses flags that may sit on either side of one positional
|
||||
// argument (`restore -yes x.tar` and `restore x.tar -yes` both work) and
|
||||
// returns that argument.
|
||||
func parseWithArg(fs *flag.FlagSet, args []string) (string, bool) {
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return "", false
|
||||
}
|
||||
if fs.NArg() == 0 {
|
||||
return "", true
|
||||
}
|
||||
arg := fs.Arg(0)
|
||||
if err := fs.Parse(fs.Args()[1:]); err != nil {
|
||||
return "", false
|
||||
}
|
||||
if fs.NArg() > 0 {
|
||||
fmt.Fprintf(fs.Output(), "felis db: unexpected argument %q\n", fs.Arg(0))
|
||||
return "", false
|
||||
}
|
||||
return arg, true
|
||||
}
|
||||
|
||||
func dbDatabaseURL(path string) (string, error) {
|
||||
cfg, err := config.Load(path)
|
||||
if err != nil {
|
||||
return "", err
|
||||
}
|
||||
return cfg.Database.URL, nil
|
||||
}
|
||||
|
||||
func dbBackup(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
label := fs.String("label", dbbackup.LabelManual, "bundle label; daily/pre-migrate/pre-restore bundles are pruned, manual ones never")
|
||||
keep := fs.Int("keep", -1, "bundles of this label to keep (default: daily 14, pre-migrate 10, pre-restore 5, manual all)")
|
||||
stateDir := fs.String("state-dir", dbbackup.DefaultStateDir, `host state directory to bundle ("" for none)`)
|
||||
noServers := fs.Bool("no-servers", false, "leave the MinecraftServer objects out of the bundle")
|
||||
metrics := fs.String("metrics-file", "", "node-exporter textfile to rewrite on success (e.g. /var/lib/node_exporter/textfile_collector/felis_db_backup.prom)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
if fs.NArg() > 0 {
|
||||
fmt.Fprint(stderr, dbUsage)
|
||||
return 2
|
||||
}
|
||||
url, err := dbDatabaseURL(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db backup: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if *keep < 0 {
|
||||
*keep = defaultKeep[*label]
|
||||
}
|
||||
o := dbbackup.BackupOptions{
|
||||
DatabaseURL: url, Dir: *dir, Label: *label, Keep: *keep,
|
||||
StateDir: *stateDir, Version: resolvedVersion(), Log: stderr,
|
||||
MetricsFile: *metrics, Record: true,
|
||||
}
|
||||
if !*noServers {
|
||||
o.ExportServers = exportMinecraftServers
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Minute)
|
||||
defer cancel()
|
||||
path, err := dbbackup.Backup(ctx, o)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db backup: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis db backup: wrote %s\n", path)
|
||||
return 0
|
||||
}
|
||||
|
||||
// resolveBundle accepts a path, or a bare bundle name looked up in dir.
|
||||
func resolveBundle(dir, arg string) string {
|
||||
if strings.ContainsRune(arg, os.PathSeparator) {
|
||||
return arg
|
||||
}
|
||||
if _, err := os.Stat(arg); err == nil {
|
||||
return arg
|
||||
}
|
||||
return filepath.Join(dir, arg)
|
||||
}
|
||||
|
||||
func dbRestore(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
yes := fs.Bool("yes", false, "replace the database's contents (required)")
|
||||
force := fs.Bool("force", false, "restore even while other clients are connected")
|
||||
noSafety := fs.Bool("no-safety-backup", false, "skip the bundle of the current database taken first")
|
||||
stateDir := fs.String("state-dir", dbbackup.DefaultStateDir, "host state directory for the safety bundle")
|
||||
arg, ok := parseWithArg(fs, args)
|
||||
if !ok {
|
||||
return 2
|
||||
}
|
||||
if arg == "" {
|
||||
fmt.Fprint(stderr, dbUsage)
|
||||
return 2
|
||||
}
|
||||
bundle := resolveBundle(*dir, arg)
|
||||
m, err := dbbackup.Verify(bundle)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if !*yes {
|
||||
fmt.Fprintf(stderr, "felis db restore: this replaces every table in the felis database with %s (%s, taken %s, schema %d).\n",
|
||||
filepath.Base(bundle), m.Label, m.CreatedAt.Format(time.RFC3339), m.SchemaVersion)
|
||||
fmt.Fprintln(stderr, "Scale felis-api and felis-operator to 0 first, then re-run with -yes.")
|
||||
return 2
|
||||
}
|
||||
url, err := dbDatabaseURL(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 60*time.Minute)
|
||||
defer cancel()
|
||||
_, safety, err := dbbackup.Restore(ctx, dbbackup.RestoreOptions{
|
||||
DatabaseURL: url, Bundle: bundle, Dir: *dir, Force: *force, SkipSafetyBackup: *noSafety,
|
||||
Safety: dbbackup.BackupOptions{Keep: defaultKeep[dbbackup.LabelPreRestore], StateDir: *stateDir,
|
||||
Version: resolvedVersion(), ExportServers: exportMinecraftServers},
|
||||
Log: stderr,
|
||||
})
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db restore: %v\n", err)
|
||||
if errors.Is(err, dbbackup.ErrClientsConnected) {
|
||||
fmt.Fprintln(stderr, " kubectl -n felis scale deployment felis-api felis-operator --replicas=0")
|
||||
}
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis db restore: restored %s (schema %d)\n", filepath.Base(bundle), m.SchemaVersion)
|
||||
if safety != "" {
|
||||
fmt.Fprintf(stdout, " the database as it was before is in %s\n", safety)
|
||||
}
|
||||
// Nothing migrates at startup, so a control plane newer than the bundle needs
|
||||
// its migrations re-applied; rolling back to the release that wrote the bundle
|
||||
// must skip that, or the rollback is undone.
|
||||
fmt.Fprintf(stdout, " next: felis migrate up -config %s (skip it when rolling back to felis %s, which wrote this bundle)\n", *cfgPath, orUnknown(m.FelisVersion))
|
||||
fmt.Fprintln(stdout, " kubectl -n felis scale deployment felis-api felis-operator --replicas=1")
|
||||
return 0
|
||||
}
|
||||
|
||||
func dbVerify(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
|
||||
arg, ok := parseWithArg(fs, args)
|
||||
if !ok {
|
||||
return 2
|
||||
}
|
||||
if arg == "" {
|
||||
fmt.Fprint(stderr, dbUsage)
|
||||
return 2
|
||||
}
|
||||
bundle := resolveBundle(*dir, arg)
|
||||
m, err := dbbackup.Verify(bundle)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db verify: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "%s: ok\n taken %s (%s)\n felis %s\n schema %d\n %s\n",
|
||||
filepath.Base(bundle), m.CreatedAt.Format(time.RFC3339), m.Label, orUnknown(m.FelisVersion), m.SchemaVersion, orUnknown(m.PGDumpVersion))
|
||||
for _, f := range m.Files {
|
||||
if f.Link != "" {
|
||||
fmt.Fprintf(stdout, " %-40s -> %s\n", f.Name, f.Link)
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(stdout, " %-40s %d bytes\n", f.Name, f.Size)
|
||||
}
|
||||
if m.ServersError != "" {
|
||||
fmt.Fprintf(stdout, " (no MinecraftServer objects: %s)\n", m.ServersError)
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func orUnknown(s string) string {
|
||||
if s == "" {
|
||||
return "unknown"
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func dbList(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
all, err := dbbackup.List(*dir)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db list: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if len(all) == 0 {
|
||||
fmt.Fprintf(stdout, "no database backups in %s\n", *dir)
|
||||
return 0
|
||||
}
|
||||
now := time.Now()
|
||||
for _, b := range all {
|
||||
fmt.Fprintf(stdout, "%-50s %-12s %10s %s ago\n", b.Name, b.Label, humanBytes(b.Size), dbbackup.Age(now.Sub(b.Created)))
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func humanBytes(n int64) string {
|
||||
const unit = 1024
|
||||
if n < unit {
|
||||
return fmt.Sprintf("%d B", n)
|
||||
}
|
||||
div, exp := int64(unit), 0
|
||||
for m := n / unit; m >= unit; m /= unit {
|
||||
div *= unit
|
||||
exp++
|
||||
}
|
||||
return fmt.Sprintf("%.1f %ciB", float64(n)/float64(div), "KMGTPE"[exp])
|
||||
}
|
||||
|
||||
// dbCheck is the freshness probe: exit 1 when the newest bundle is missing or
|
||||
// older than -max-age, for a monitor or the break-glass console to act on.
|
||||
func dbCheck(fs *flag.FlagSet, dir *string, args []string, stdout, stderr io.Writer) int {
|
||||
maxAge := fs.Duration("max-age", dbbackup.StaleAfter, "oldest acceptable newest bundle")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
b, err := dbbackup.Check(*dir, *maxAge, time.Now())
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis db check: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis db check: ok, newest backup %s (%s ago)\n", b.Name, dbbackup.Age(time.Since(b.Created)))
|
||||
return 0
|
||||
}
|
||||
|
||||
// exportMinecraftServers reads every MinecraftServer through the host's k3s
|
||||
// kubectl and strips what the API server owns, so the result can be fed back
|
||||
// with `kubectl apply -f` on a rebuilt cluster.
|
||||
func exportMinecraftServers(ctx context.Context) ([]byte, error) {
|
||||
ctx, cancel := context.WithTimeout(ctx, 30*time.Second)
|
||||
defer cancel()
|
||||
// Output, not the CombinedOutput kubectlOutput uses: a deprecation warning
|
||||
// on stderr must not end up inside the JSON.
|
||||
cmd := exec.CommandContext(ctx, "k3s", "kubectl", "get", "minecraftservers.felis.lolicon.best", "-A", "-o", "json")
|
||||
cmd.Env = append(os.Environ(), "KUBECONFIG="+hostBootstrapKubeconfigPath)
|
||||
var errBuf strings.Builder
|
||||
cmd.Stderr = &errBuf
|
||||
out, err := cmd.Output()
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("k3s kubectl get minecraftservers: %w: %s", err, strings.TrimSpace(errBuf.String()))
|
||||
}
|
||||
return cleanServerList(out)
|
||||
}
|
||||
|
||||
// cleanServerList drops status and the server-assigned metadata from a
|
||||
// `kubectl get -o json` List.
|
||||
func cleanServerList(raw []byte) ([]byte, error) {
|
||||
var list struct {
|
||||
Items []map[string]any `json:"items"`
|
||||
}
|
||||
if err := json.Unmarshal(raw, &list); err != nil {
|
||||
return nil, fmt.Errorf("parse MinecraftServer list: %w", err)
|
||||
}
|
||||
for _, it := range list.Items {
|
||||
delete(it, "status")
|
||||
if md, ok := it["metadata"].(map[string]any); ok {
|
||||
for _, k := range []string{"resourceVersion", "uid", "creationTimestamp", "generation", "managedFields", "selfLink"} {
|
||||
delete(md, k)
|
||||
}
|
||||
}
|
||||
}
|
||||
if list.Items == nil {
|
||||
list.Items = []map[string]any{}
|
||||
}
|
||||
return json.MarshalIndent(map[string]any{"apiVersion": "v1", "kind": "List", "items": list.Items}, "", " ")
|
||||
}
|
||||
@@ -0,0 +1,151 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"flag"
|
||||
"io"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/store"
|
||||
)
|
||||
|
||||
func TestDBUsage(t *testing.T) {
|
||||
for _, args := range [][]string{{"db"}, {"db", "frobnicate"}, {"db", "restore"}, {"db", "verify"}, {"db", "backup", "extra"}} {
|
||||
var out, errBuf bytes.Buffer
|
||||
if code := run(args, &out, &errBuf); code != 2 {
|
||||
t.Errorf("%v: exit %d, want 2", args, code)
|
||||
}
|
||||
if !strings.Contains(errBuf.String(), "felis db restore") {
|
||||
t.Errorf("%v: no usage on stderr: %q", args, errBuf.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestDBRestoreNeedsYes(t *testing.T) {
|
||||
// A bundle that does not exist fails verification (1) before -yes matters;
|
||||
// the -yes gate itself is exercised against a real bundle in internal/dbbackup
|
||||
// and on the VM. Here: the refusal path never reaches the config or database.
|
||||
var out, errBuf bytes.Buffer
|
||||
if code := run([]string{"db", "restore", "-dir", t.TempDir(), "missing.tar"}, &out, &errBuf); code != 1 {
|
||||
t.Fatalf("exit %d, stderr %q", code, errBuf.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseWithArg(t *testing.T) {
|
||||
for _, args := range [][]string{{"-yes", "b.tar"}, {"b.tar", "-yes"}} {
|
||||
fs := flag.NewFlagSet("t", flag.ContinueOnError)
|
||||
fs.SetOutput(io.Discard)
|
||||
yes := fs.Bool("yes", false, "")
|
||||
arg, ok := parseWithArg(fs, args)
|
||||
if !ok || arg != "b.tar" || !*yes {
|
||||
t.Errorf("%v -> %q ok=%v yes=%v", args, arg, ok, *yes)
|
||||
}
|
||||
}
|
||||
fs := flag.NewFlagSet("t", flag.ContinueOnError)
|
||||
fs.SetOutput(io.Discard)
|
||||
if _, ok := parseWithArg(fs, []string{"a.tar", "b.tar"}); ok {
|
||||
t.Error("two positional arguments accepted")
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveBundle(t *testing.T) {
|
||||
if got := resolveBundle("/var/lib/felis/db-backups", "felis-db-x.tar"); got != "/var/lib/felis/db-backups/felis-db-x.tar" {
|
||||
t.Errorf("bare name -> %s", got)
|
||||
}
|
||||
if got := resolveBundle("/var/lib/felis/db-backups", "/root/copy.tar"); got != "/root/copy.tar" {
|
||||
t.Errorf("path -> %s", got)
|
||||
}
|
||||
}
|
||||
|
||||
func TestCleanServerList(t *testing.T) {
|
||||
raw := `{"apiVersion":"v1","kind":"List","metadata":{"resourceVersion":""},"items":[{
|
||||
"apiVersion":"felis.lolicon.best/v1alpha1","kind":"MinecraftServer",
|
||||
"metadata":{"name":"survival","namespace":"minecraft","uid":"u","resourceVersion":"42","generation":3,
|
||||
"creationTimestamp":"2026-09-01T00:00:00Z","managedFields":[{}],"labels":{"a":"b"}},
|
||||
"spec":{"desiredState":"Running"},"status":{"phase":"Running"}}]}`
|
||||
out, err := cleanServerList([]byte(raw))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var got struct {
|
||||
Kind string `json:"kind"`
|
||||
Items []map[string]any `json:"items"`
|
||||
}
|
||||
if err := json.Unmarshal(out, &got); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if got.Kind != "List" || len(got.Items) != 1 {
|
||||
t.Fatalf("got %s", out)
|
||||
}
|
||||
it := got.Items[0]
|
||||
if _, ok := it["status"]; ok {
|
||||
t.Error("status kept")
|
||||
}
|
||||
md := it["metadata"].(map[string]any)
|
||||
for _, k := range []string{"uid", "resourceVersion", "generation", "creationTimestamp", "managedFields"} {
|
||||
if _, ok := md[k]; ok {
|
||||
t.Errorf("metadata.%s kept", k)
|
||||
}
|
||||
}
|
||||
if md["name"] != "survival" || md["namespace"] != "minecraft" || md["labels"] == nil {
|
||||
t.Errorf("identity lost: %v", md)
|
||||
}
|
||||
if it["spec"].(map[string]any)["desiredState"] != "Running" {
|
||||
t.Error("spec lost")
|
||||
}
|
||||
|
||||
empty, err := cleanServerList([]byte(`{"items":null}`))
|
||||
if err != nil || !strings.Contains(string(empty), `"items": []`) {
|
||||
t.Errorf("empty list -> %s, %v", empty, err)
|
||||
}
|
||||
if _, err := cleanServerList([]byte("Warning: x\n{")); err == nil {
|
||||
t.Error("garbage parsed")
|
||||
}
|
||||
}
|
||||
|
||||
func TestHasPending(t *testing.T) {
|
||||
ms := []store.Migration{{Version: 1}, {Version: 2}, {Version: 3}}
|
||||
if hasPending(map[int]struct{}{1: {}, 2: {}, 3: {}}, ms) {
|
||||
t.Error("fully applied reported pending")
|
||||
}
|
||||
if !hasPending(map[int]struct{}{1: {}, 2: {}}, ms) {
|
||||
t.Error("missing 3 not reported")
|
||||
}
|
||||
}
|
||||
|
||||
type appliedDriver struct {
|
||||
store.Driver
|
||||
done map[int]struct{}
|
||||
}
|
||||
|
||||
func (d appliedDriver) EnsureVersionTable(context.Context) error { return nil }
|
||||
func (d appliedDriver) AppliedVersions(context.Context) (map[int]struct{}, error) {
|
||||
return d.done, nil
|
||||
}
|
||||
|
||||
func TestPreMigrateBackupOnlyGuardsAPopulatedDatabase(t *testing.T) {
|
||||
ms := []store.Migration{{Version: 1}, {Version: 2}}
|
||||
// An unusable URL makes an attempted backup observable as an error without
|
||||
// any PostgreSQL tooling.
|
||||
const badURL = "not-a-url"
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
done map[int]struct{}
|
||||
attempt bool
|
||||
}{
|
||||
{"fresh database", map[int]struct{}{}, false},
|
||||
{"up to date", map[int]struct{}{1: {}, 2: {}}, false},
|
||||
{"pending on a populated database", map[int]struct{}{1: {}}, true},
|
||||
} {
|
||||
path, err := preMigrateBackup(context.Background(), appliedDriver{done: tc.done}, ms, badURL, t.TempDir(), io.Discard)
|
||||
if attempted := err != nil; attempted != tc.attempt {
|
||||
t.Errorf("%s: attempted = %v (err %v), want %v", tc.name, attempted, err, tc.attempt)
|
||||
}
|
||||
if path != "" {
|
||||
t.Errorf("%s: path = %q", tc.name, path)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,66 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"net"
|
||||
"os"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Vars so tests can shrink them. A dial that neither connects nor is refused
|
||||
// within egressDialTimeout counts as blocked: a policy that drops packets looks
|
||||
// exactly like that.
|
||||
var (
|
||||
egressDialTimeout = 500 * time.Millisecond
|
||||
egressPollInterval = 200 * time.Millisecond
|
||||
)
|
||||
|
||||
// cmdEgressGate is the first initContainer of every build pod. The pod's
|
||||
// NetworkPolicy is programmed asynchronously after the pod starts (live on k3s:
|
||||
// a build-labelled pod reached the internet and the Kubernetes API for its first
|
||||
// ~0.7 s), so the gate dials a destination the policy denies until it stops
|
||||
// answering, and only then lets the pod's next container, eventually the
|
||||
// untrusted Dockerfile, start.
|
||||
//
|
||||
// The default probe is the Kubernetes API Service, which the kubelet names in
|
||||
// every pod's environment and the build policy never admits. A probe that still
|
||||
// answers after --wait means the policy is not enforced at all (a CNI without
|
||||
// NetworkPolicy support, or k3s run with --disable-network-policy), and the
|
||||
// build fails closed.
|
||||
func cmdEgressGate(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("egress-gate", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
probe := fs.String("probe", "", "host:port the build NetworkPolicy denies (default: the Kubernetes API Service from KUBERNETES_SERVICE_HOST/PORT)")
|
||||
wait := fs.Duration("wait", 2*time.Minute, "how long the probe may keep answering before the build is refused")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
if *probe == "" {
|
||||
host, port := os.Getenv("KUBERNETES_SERVICE_HOST"), os.Getenv("KUBERNETES_SERVICE_PORT")
|
||||
if host == "" || port == "" {
|
||||
fmt.Fprintln(stderr, "felis egress-gate: no --probe and no KUBERNETES_SERVICE_HOST/PORT to default to")
|
||||
return 2
|
||||
}
|
||||
*probe = net.JoinHostPort(host, port)
|
||||
}
|
||||
|
||||
start := time.Now()
|
||||
for {
|
||||
conn, err := net.DialTimeout("tcp", *probe, egressDialTimeout)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stdout, "felis egress-gate: %s is unreachable after %s (%v); the egress lock is in effect\n",
|
||||
*probe, time.Since(start).Round(time.Millisecond), err)
|
||||
return 0
|
||||
}
|
||||
_ = conn.Close()
|
||||
if time.Since(start) >= *wait {
|
||||
fmt.Fprintf(stderr, "felis egress-gate: %s still answers after %s: the build namespace's NetworkPolicy is not enforced "+
|
||||
"(a CNI without NetworkPolicy support, or k3s started with --disable-network-policy); refusing to run the build\n",
|
||||
*probe, *wait)
|
||||
return 1
|
||||
}
|
||||
time.Sleep(egressPollInterval)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,98 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"net"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
func shrinkEgressGate(t *testing.T) {
|
||||
t.Helper()
|
||||
dial, poll := egressDialTimeout, egressPollInterval
|
||||
egressDialTimeout, egressPollInterval = 200*time.Millisecond, 10*time.Millisecond
|
||||
t.Cleanup(func() { egressDialTimeout, egressPollInterval = dial, poll })
|
||||
}
|
||||
|
||||
// The gate holds while the probe answers and lets the pod go on once the policy
|
||||
// lands, which the test plays by closing the listener.
|
||||
func TestEgressGateWaitsForTheLock(t *testing.T) {
|
||||
shrinkEgressGate(t)
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
accepted := make(chan struct{}, 100)
|
||||
go func() {
|
||||
for {
|
||||
c, err := ln.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
_ = c.Close()
|
||||
accepted <- struct{}{}
|
||||
}
|
||||
}()
|
||||
go func() {
|
||||
for i := 0; i < 3; i++ {
|
||||
<-accepted
|
||||
}
|
||||
_ = ln.Close()
|
||||
}()
|
||||
var out, errb bytes.Buffer
|
||||
if code := cmdEgressGate([]string{"--probe", ln.Addr().String(), "--wait", "10s"}, &out, &errb); code != 0 {
|
||||
t.Fatalf("exit %d: %s", code, errb.String())
|
||||
}
|
||||
if !strings.Contains(out.String(), "egress lock is in effect") {
|
||||
t.Errorf("stdout = %q", out.String())
|
||||
}
|
||||
}
|
||||
|
||||
// A probe that keeps answering means no policy is enforced: the build must not run.
|
||||
func TestEgressGateRefusesAnOpenNetwork(t *testing.T) {
|
||||
shrinkEgressGate(t)
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
defer ln.Close()
|
||||
go func() {
|
||||
for {
|
||||
c, err := ln.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
_ = c.Close()
|
||||
}
|
||||
}()
|
||||
var out, errb bytes.Buffer
|
||||
if code := cmdEgressGate([]string{"--probe", ln.Addr().String(), "--wait", "100ms"}, &out, &errb); code != 1 {
|
||||
t.Fatalf("exit %d, want 1", code)
|
||||
}
|
||||
if !strings.Contains(errb.String(), "not enforced") {
|
||||
t.Errorf("stderr = %q", errb.String())
|
||||
}
|
||||
}
|
||||
|
||||
func TestEgressGateDefaultsToTheKubernetesService(t *testing.T) {
|
||||
shrinkEgressGate(t)
|
||||
t.Setenv("KUBERNETES_SERVICE_HOST", "")
|
||||
t.Setenv("KUBERNETES_SERVICE_PORT", "")
|
||||
var out, errb bytes.Buffer
|
||||
if code := cmdEgressGate(nil, &out, &errb); code != 2 {
|
||||
t.Fatalf("exit %d without a probe, want 2", code)
|
||||
}
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
host, port, _ := net.SplitHostPort(ln.Addr().String())
|
||||
_ = ln.Close() // closed: the lock reads as in effect at once
|
||||
t.Setenv("KUBERNETES_SERVICE_HOST", host)
|
||||
t.Setenv("KUBERNETES_SERVICE_PORT", port)
|
||||
out.Reset()
|
||||
if code := cmdEgressGate(nil, &out, &errb); code != 0 || !strings.Contains(out.String(), ln.Addr().String()) {
|
||||
t.Fatalf("exit %d, stdout %q", code, out.String())
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,259 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"archive/tar"
|
||||
"compress/gzip"
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"net/http"
|
||||
"os"
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/build"
|
||||
)
|
||||
|
||||
// cmdFetchContext is the in-Pod entrypoint the build Job's context-fetch
|
||||
// initContainer runs. It reads the blob the platform stored for a submission
|
||||
// from the felis-api INTERNAL face (with a bounded retry — see
|
||||
// fetchContextWithRetry) and extracts it into the shared emptyDir the Kaniko
|
||||
// container then builds from.
|
||||
//
|
||||
// Why this exists: the build Pod runs in the build namespace, where it can neither
|
||||
// mount the control-plane uploads PVC (a PVC does not cross namespaces) nor hold
|
||||
// object-store credentials, so the API that WROTE the blob is the transport. The
|
||||
// route is service-token-gated; the token arrives through a namespace-local Secret
|
||||
// mounted only into this initContainer, never into Kaniko's — so the untrusted
|
||||
// Dockerfile's build steps have no credential to read (their containers share no
|
||||
// environment, no PID namespace, and Kaniko itself mounts the context read-only).
|
||||
//
|
||||
// The extraction is deliberately paranoid: the tarball is attacker-controlled
|
||||
// input, so absolute paths, ".." escapes, links, and special files are refused
|
||||
// rather than sanitized. Kaniko treats the extracted tree as hostile regardless
|
||||
// (spec §16), but the pod's own filesystem still must not be written outside the
|
||||
// context directory it was given.
|
||||
func cmdFetchContext(args []string, _, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("fetch-context", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
url := fs.String("url", "", "internal-face URL of the submission's build-context tarball")
|
||||
out := fs.String("out", "/context", "directory to extract the build context into")
|
||||
want := fs.String("sha256", "", "refuse the context unless the tarball's sha256 is this lowercase hex digest")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
if *want != "" && !build.IsSHA256Hex(*want) {
|
||||
fmt.Fprintf(stderr, "felis fetch-context: --sha256 %q is not a lowercase hex sha256\n", *want)
|
||||
return 2
|
||||
}
|
||||
if *url == "" {
|
||||
fmt.Fprintln(stderr, "felis fetch-context: --url is required")
|
||||
return 2
|
||||
}
|
||||
token := os.Getenv("FELIS_SERVICE_TOKEN")
|
||||
if token == "" {
|
||||
fmt.Fprintln(stderr, "felis fetch-context: FELIS_SERVICE_TOKEN is empty — the internal face rejects anonymous reads")
|
||||
return 2
|
||||
}
|
||||
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
|
||||
// Validate the URL once up front: a bad one is a usage error (2), not
|
||||
// something to sit in the retry loop.
|
||||
if _, err := http.NewRequest(http.MethodGet, *url, nil); err != nil {
|
||||
fmt.Fprintf(stderr, "felis fetch-context: bad --url: %v\n", err)
|
||||
return 2
|
||||
}
|
||||
// No overall client timeout: a legitimate modpack context can be large and the
|
||||
// Job's activeDeadlineSeconds is the real bound. The header timeout catches a
|
||||
// wedged endpoint without capping a healthy download.
|
||||
// Redirects are refused: the request carries the service token, and the
|
||||
// internal face never redirects, so a 3xx is someone steering the token.
|
||||
client := &http.Client{
|
||||
Transport: &http.Transport{ResponseHeaderTimeout: time.Minute},
|
||||
CheckRedirect: func(*http.Request, []*http.Request) error { return http.ErrUseLastResponse },
|
||||
}
|
||||
resp, err := fetchContextWithRetry(ctx, client, *url, token, stderr)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer resp.Body.Close()
|
||||
|
||||
h := sha256.New()
|
||||
body := io.TeeReader(resp.Body, h)
|
||||
if err := extractTarGz(body, *out); err != nil {
|
||||
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if *want == "" {
|
||||
return 0
|
||||
}
|
||||
// The tar end marker comes before the gzip trailer and whatever follows it,
|
||||
// so read to EOF: the digest must cover every byte the blob holds. The blob
|
||||
// itself is size-capped at upload, which bounds this read.
|
||||
if _, err := io.Copy(io.Discard, io.LimitReader(body, maxContextBytes)); err != nil {
|
||||
fmt.Fprintf(stderr, "felis fetch-context: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if got := hex.EncodeToString(h.Sum(nil)); got != *want {
|
||||
// The init container failing is what keeps Kaniko from ever starting on
|
||||
// the extracted tree.
|
||||
fmt.Fprintf(stderr, "felis fetch-context: the context's sha256 is %s, the approved digest is %s: it changed after approval; refusing to build\n", got, *want)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// fetchRetryInterval/fetchRetryWindow bound how long the fetch waits out a
|
||||
// control-plane blip before giving up. The api pod being replaced is a normal
|
||||
// event (rollout, eviction, a chaos drill), and without a retry one refused
|
||||
// dial turns it into a failed build: BackoffLimit=0 gives the Job no second
|
||||
// Pod, so the terminal verdict costs a manual re-approval — the live drill hit
|
||||
// exactly this (context-fetch exit 1 on `connect: connection refused` while
|
||||
// the api pod rolled; the new pod was serving 11 seconds later and the same
|
||||
// 198-byte blob). The window is tiny next to the Job's 30-minute
|
||||
// activeDeadline; a 4xx (missing blob, rejected token) still fails fast.
|
||||
//
|
||||
// Vars, not consts, so tests can shrink the window.
|
||||
var (
|
||||
fetchRetryInterval = 3 * time.Second
|
||||
fetchRetryWindow = 45 * time.Second
|
||||
)
|
||||
|
||||
// fetchContextWithRetry GETs the context tarball, retrying transport failures
|
||||
// and 5xx responses until fetchRetryWindow runs out. A 4xx is an answer, not a
|
||||
// blip — retrying it only delays the honest error.
|
||||
func fetchContextWithRetry(ctx context.Context, client *http.Client, url, token string, stderr io.Writer) (*http.Response, error) {
|
||||
deadline := time.Now().Add(fetchRetryWindow)
|
||||
for {
|
||||
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("bad --url: %w", err)
|
||||
}
|
||||
req.Header.Set("Authorization", "Bearer "+token)
|
||||
|
||||
resp, err := client.Do(req)
|
||||
if err == nil && resp.StatusCode == http.StatusOK {
|
||||
return resp, nil
|
||||
}
|
||||
if err == nil {
|
||||
status := resp.Status
|
||||
_ = resp.Body.Close()
|
||||
err = fmt.Errorf("GET returned %s", status)
|
||||
if resp.StatusCode < 500 {
|
||||
return nil, err
|
||||
}
|
||||
}
|
||||
if ctx.Err() != nil {
|
||||
return nil, fmt.Errorf("GET failed: %w", err)
|
||||
}
|
||||
if time.Now().After(deadline) {
|
||||
return nil, fmt.Errorf("GET failed (retried for %s): %w", fetchRetryWindow, err)
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis fetch-context: %v; retrying (the internal face may be restarting)\n", err)
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return nil, fmt.Errorf("GET failed: %w", err)
|
||||
case <-time.After(fetchRetryInterval):
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// maxContextBytes / maxContextEntries bound what one context may expand to. The
|
||||
// compressed upload is capped at 1 GiB, but gzip turns that into hundreds of GiB
|
||||
// or millions of empty files, and the emptyDir's 4 GiB sizeLimit is only
|
||||
// enforced by the kubelet's periodic sweep, after the disk has filled. The byte
|
||||
// cap matches that sizeLimit; the entry cap is far above any real modpack (a
|
||||
// large one is a few thousand files) and far below an inode exhaustion.
|
||||
//
|
||||
// Vars, not consts, so tests can shrink them.
|
||||
var (
|
||||
maxContextBytes int64 = 4 << 30
|
||||
maxContextEntries = 200_000
|
||||
)
|
||||
|
||||
// extractTarGz streams a gzip'd tarball into root, creating directories as
|
||||
// needed. Every entry is vetted BEFORE anything is written: a path that is
|
||||
// absolute or escapes root (via ".."), a link (symlink or hardlink), or any
|
||||
// special file kind aborts the whole extraction. Refusing rather than skipping is
|
||||
// deliberate — a context that needs one of those constructs is not a context this
|
||||
// transport carries, and silently dropping entries would build from a corpus the
|
||||
// submitter did not upload. The whole extraction is also bounded by
|
||||
// maxContextBytes and maxContextEntries.
|
||||
func extractTarGz(r io.Reader, root string) error {
|
||||
if err := os.MkdirAll(root, 0o755); err != nil {
|
||||
return fmt.Errorf("create context dir: %w", err)
|
||||
}
|
||||
zr, err := gzip.NewReader(r)
|
||||
if err != nil {
|
||||
return fmt.Errorf("context is not a valid gzip tarball: %w", err)
|
||||
}
|
||||
defer zr.Close()
|
||||
tr := tar.NewReader(zr)
|
||||
var written int64
|
||||
entries := 0
|
||||
for {
|
||||
hdr, err := tr.Next()
|
||||
if errors.Is(err, io.EOF) {
|
||||
return nil
|
||||
}
|
||||
if err != nil {
|
||||
return fmt.Errorf("read context tarball: %w", err)
|
||||
}
|
||||
if entries++; entries > maxContextEntries {
|
||||
return fmt.Errorf("the build context has more than %d entries", maxContextEntries)
|
||||
}
|
||||
name := filepath.Clean(hdr.Name)
|
||||
if name == "." {
|
||||
continue
|
||||
}
|
||||
// The zip-slip guard: reject, never rewrite. filepath.Clean collapses any
|
||||
// "a/../../b", so these two checks are sufficient once Clean has run.
|
||||
if filepath.IsAbs(name) || name == ".." || strings.HasPrefix(name, ".."+string(filepath.Separator)) {
|
||||
return fmt.Errorf("context entry %q escapes the context directory", hdr.Name)
|
||||
}
|
||||
target := filepath.Join(root, name)
|
||||
switch hdr.Typeflag {
|
||||
case tar.TypeDir:
|
||||
if err := os.MkdirAll(target, 0o755); err != nil {
|
||||
return fmt.Errorf("create %q: %w", name, err)
|
||||
}
|
||||
case tar.TypeReg:
|
||||
if err := os.MkdirAll(filepath.Dir(target), 0o755); err != nil {
|
||||
return fmt.Errorf("create parent of %q: %w", name, err)
|
||||
}
|
||||
mode := os.FileMode(0o644)
|
||||
if hdr.FileInfo().Mode()&0o111 != 0 {
|
||||
mode = 0o755 // preserve executability (entrypoint scripts), nothing else
|
||||
}
|
||||
f, err := os.OpenFile(target, os.O_CREATE|os.O_WRONLY|os.O_TRUNC, mode)
|
||||
if err != nil {
|
||||
return fmt.Errorf("create %q: %w", name, err)
|
||||
}
|
||||
n, err := io.Copy(f, io.LimitReader(tr, maxContextBytes-written+1))
|
||||
written += n
|
||||
if err != nil {
|
||||
_ = f.Close()
|
||||
return fmt.Errorf("write %q: %w", name, err)
|
||||
}
|
||||
if written > maxContextBytes {
|
||||
_ = f.Close()
|
||||
return fmt.Errorf("the build context expands past %d bytes", maxContextBytes)
|
||||
}
|
||||
if err := f.Close(); err != nil {
|
||||
return fmt.Errorf("close %q: %w", name, err)
|
||||
}
|
||||
default:
|
||||
return fmt.Errorf("context entry %q has unsupported type %q (links and special files are refused)", hdr.Name, string(hdr.Typeflag))
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,404 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"archive/tar"
|
||||
"bytes"
|
||||
"compress/gzip"
|
||||
"crypto/sha256"
|
||||
"encoding/hex"
|
||||
"io"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
"time"
|
||||
)
|
||||
|
||||
type tarEntry struct {
|
||||
name string
|
||||
body string
|
||||
mode int64
|
||||
typ byte
|
||||
linkname string
|
||||
}
|
||||
|
||||
// tgzBody builds an in-memory .tar.gz from entries, preserving each entry's type
|
||||
// and mode so the tests can exercise the guards with exactly the bytes an
|
||||
// attacker could upload.
|
||||
func tgzBody(t *testing.T, entries ...tarEntry) []byte {
|
||||
t.Helper()
|
||||
var buf bytes.Buffer
|
||||
zw := gzip.NewWriter(&buf)
|
||||
tw := tar.NewWriter(zw)
|
||||
for _, e := range entries {
|
||||
typ := e.typ
|
||||
if typ == 0 {
|
||||
typ = tar.TypeReg
|
||||
}
|
||||
mode := e.mode
|
||||
if mode == 0 {
|
||||
mode = 0o644
|
||||
}
|
||||
hdr := &tar.Header{Name: e.name, Typeflag: typ, Mode: mode, Size: int64(len(e.body))}
|
||||
if typ == tar.TypeSymlink {
|
||||
hdr.Linkname = e.linkname
|
||||
hdr.Size = 0
|
||||
}
|
||||
if err := tw.WriteHeader(hdr); err != nil {
|
||||
t.Fatalf("write header %q: %v", e.name, err)
|
||||
}
|
||||
if hdr.Size > 0 {
|
||||
if _, err := tw.Write([]byte(e.body)); err != nil {
|
||||
t.Fatalf("write body %q: %v", e.name, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
if err := tw.Close(); err != nil {
|
||||
t.Fatalf("close tar: %v", err)
|
||||
}
|
||||
if err := zw.Close(); err != nil {
|
||||
t.Fatalf("close gzip: %v", err)
|
||||
}
|
||||
return buf.Bytes()
|
||||
}
|
||||
|
||||
// A normal context extracts with its tree intact, and the executable bit that
|
||||
// modpack entrypoints rely on survives.
|
||||
func TestExtractTarGzRoundTrip(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
body := tgzBody(t,
|
||||
tarEntry{name: "Dockerfile", body: "FROM scratch\n"},
|
||||
tarEntry{name: "mods/example.jar", body: "jar-bytes"},
|
||||
tarEntry{name: "start.sh", body: "#!/bin/sh\n", mode: 0o755},
|
||||
tarEntry{name: "mods/", typ: tar.TypeDir, mode: 0o755},
|
||||
)
|
||||
if err := extractTarGz(bytes.NewReader(body), dir); err != nil {
|
||||
t.Fatalf("extract: %v", err)
|
||||
}
|
||||
for name, want := range map[string]string{
|
||||
"Dockerfile": "FROM scratch\n",
|
||||
"mods/example.jar": "jar-bytes",
|
||||
} {
|
||||
got, err := os.ReadFile(filepath.Join(dir, name))
|
||||
if err != nil || string(got) != want {
|
||||
t.Fatalf("%s = (%q, %v), want %q", name, got, err, want)
|
||||
}
|
||||
}
|
||||
fi, err := os.Stat(filepath.Join(dir, "start.sh"))
|
||||
if err != nil || fi.Mode()&0o111 == 0 {
|
||||
t.Fatalf("entrypoint script lost its exec bit: %v (%v)", fi, err)
|
||||
}
|
||||
}
|
||||
|
||||
// The guards: "..", absolute paths, symlinks, and special files are refused whole
|
||||
// — nothing escapes, and nothing is silently skipped.
|
||||
func TestExtractTarGzRefusesEscapes(t *testing.T) {
|
||||
cases := []struct {
|
||||
name string
|
||||
entries []tarEntry
|
||||
}{
|
||||
{"dotdot", []tarEntry{{name: "../outside", body: "x"}}},
|
||||
{"nested dotdot", []tarEntry{{name: "a/../../outside", body: "x"}}},
|
||||
{"absolute", []tarEntry{{name: "/etc/outside", body: "x"}}},
|
||||
{"symlink", []tarEntry{{name: "link", typ: tar.TypeSymlink, linkname: "/etc"}}},
|
||||
{"hardlink", []tarEntry{{name: "hard", typ: tar.TypeLink, linkname: "somewhere"}}},
|
||||
{"device", []tarEntry{{name: "dev", typ: tar.TypeChar}}},
|
||||
}
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
if err := extractTarGz(bytes.NewReader(tgzBody(t, tc.entries...)), dir); err == nil {
|
||||
t.Fatal("extract accepted a hostile entry, want an error")
|
||||
}
|
||||
// Nothing may have been written outside the target (or at all).
|
||||
entries, _ := os.ReadDir(dir)
|
||||
if len(entries) != 0 {
|
||||
t.Fatalf("hostile archive left %d entries behind", len(entries))
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// A context that expands past the byte or entry cap is refused, however small
|
||||
// it was compressed: gzip bombs and inode floods stop at the cap.
|
||||
func TestExtractTarGzCapsExpansion(t *testing.T) {
|
||||
bytesCap, entriesCap := maxContextBytes, maxContextEntries
|
||||
t.Cleanup(func() { maxContextBytes, maxContextEntries = bytesCap, entriesCap })
|
||||
maxContextBytes, maxContextEntries = 1000, 5
|
||||
|
||||
fits := tgzBody(t, tarEntry{name: "a", body: strings.Repeat("x", 600)}, tarEntry{name: "b", body: strings.Repeat("y", 400)})
|
||||
if err := extractTarGz(bytes.NewReader(fits), t.TempDir()); err != nil {
|
||||
t.Fatalf("a context exactly at the byte cap: %v", err)
|
||||
}
|
||||
big := tgzBody(t, tarEntry{name: "a", body: strings.Repeat("x", 600)}, tarEntry{name: "b", body: strings.Repeat("y", 401)})
|
||||
if err := extractTarGz(bytes.NewReader(big), t.TempDir()); err == nil || !strings.Contains(err.Error(), "expands past") {
|
||||
t.Fatalf("one byte over the cap: err = %v", err)
|
||||
}
|
||||
var many []tarEntry
|
||||
for i := 0; i < 6; i++ {
|
||||
many = append(many, tarEntry{name: "d" + string(rune('0'+i)) + "/", typ: tar.TypeDir})
|
||||
}
|
||||
if err := extractTarGz(bytes.NewReader(tgzBody(t, many...)), t.TempDir()); err == nil || !strings.Contains(err.Error(), "entries") {
|
||||
t.Fatalf("six entries over a cap of five: err = %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
// The command end to end: it dials the URL with the bearer token from the
|
||||
// environment, and refuses to run without it (the internal face would 401
|
||||
// anyway; failing at parse time is the honest earlier error).
|
||||
func TestCmdFetchContextFetchAndExtract(t *testing.T) {
|
||||
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.Header.Get("Authorization") != "Bearer test-token" {
|
||||
w.WriteHeader(http.StatusUnauthorized)
|
||||
return
|
||||
}
|
||||
w.Header().Set("Content-Type", "application/gzip")
|
||||
_, _ = w.Write(body)
|
||||
}))
|
||||
defer srv.Close()
|
||||
|
||||
dir := t.TempDir()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, io.Discard); code != 0 {
|
||||
t.Fatalf("cmdFetchContext exit = %d, want 0", code)
|
||||
}
|
||||
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
|
||||
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
|
||||
}
|
||||
|
||||
// No token: refuse before dialing.
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "")
|
||||
var stderr bytes.Buffer
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 2 {
|
||||
t.Fatalf("missing token exit = %d, want 2 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
|
||||
// A non-200 answer (e.g. the route's 404 for a never-uploaded context) fails.
|
||||
srv404 := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
w.WriteHeader(http.StatusNotFound)
|
||||
}))
|
||||
defer srv404.Close()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
if code := cmdFetchContext([]string{"--url=" + srv404.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, io.Discard); code != 1 {
|
||||
t.Fatalf("404 exit = %d, want 1", code)
|
||||
}
|
||||
}
|
||||
|
||||
// With --sha256 the fetch refuses any bytes but the approved ones, including
|
||||
// a tarball that extracts cleanly: that is exactly the context an uploader
|
||||
// swapped in after the review.
|
||||
func TestCmdFetchContextChecksDigest(t *testing.T) {
|
||||
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
_, _ = w.Write(body)
|
||||
}))
|
||||
defer srv.Close()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
sum := sha256.Sum256(body)
|
||||
good := hex.EncodeToString(sum[:])
|
||||
args := func(digest string) []string {
|
||||
return []string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir(), "--sha256=" + digest}
|
||||
}
|
||||
|
||||
if code := cmdFetchContext(args(good), io.Discard, io.Discard); code != 0 {
|
||||
t.Fatalf("matching digest exit = %d, want 0", code)
|
||||
}
|
||||
var stderr bytes.Buffer
|
||||
other := strings.Repeat("0", 64)
|
||||
if code := cmdFetchContext(args(other), io.Discard, &stderr); code != 1 || !strings.Contains(stderr.String(), "changed after approval") {
|
||||
t.Fatalf("mismatched digest exit = %d, stderr %q; want 1 naming the change", code, stderr.String())
|
||||
}
|
||||
stderr.Reset()
|
||||
if code := cmdFetchContext(args("ABC"), io.Discard, &stderr); code != 2 {
|
||||
t.Fatalf("malformed digest exit = %d, want 2 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
|
||||
// Bytes after the tar end marker still count: appending to an approved blob
|
||||
// must change what the fetch accepts.
|
||||
padded := append(append([]byte{}, body...), "trailing"...)
|
||||
srvPadded := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
_, _ = w.Write(padded)
|
||||
}))
|
||||
defer srvPadded.Close()
|
||||
if code := cmdFetchContext([]string{"--url=" + srvPadded.URL + "/c", "--out=" + t.TempDir(), "--sha256=" + good}, io.Discard, io.Discard); code != 1 {
|
||||
t.Fatalf("padded blob exit = %d, want 1", code)
|
||||
}
|
||||
}
|
||||
|
||||
// The request carries the service token, so a redirect is a failure: the token
|
||||
// never follows it to another host (build-supply-chain-13).
|
||||
func TestCmdFetchContextRefusesRedirects(t *testing.T) {
|
||||
var leaked bool
|
||||
elsewhere := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
leaked = true
|
||||
}))
|
||||
defer elsewhere.Close()
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
http.Redirect(w, r, elsewhere.URL+"/steal", http.StatusFound)
|
||||
}))
|
||||
defer srv.Close()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
var stderr bytes.Buffer
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/c", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
|
||||
t.Fatalf("redirect exit = %d, want 1 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
if leaked {
|
||||
t.Fatal("the fetch followed the redirect")
|
||||
}
|
||||
}
|
||||
|
||||
// A body that is not a gzip tarball must fail the extraction rather than produce
|
||||
// an empty (or partial) context Kaniko would then try to build.
|
||||
func TestExtractTarGzRejectsNonGzip(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
err := extractTarGz(strings.NewReader("not a tarball"), dir)
|
||||
if err == nil || !strings.Contains(err.Error(), "gzip") {
|
||||
t.Fatalf("err = %v, want a gzip complaint", err)
|
||||
}
|
||||
}
|
||||
|
||||
// shrinkFetchWindow swaps the retry knobs for a faster test and restores them
|
||||
// afterwards, so no test leaks a tiny window into another.
|
||||
func shrinkFetchWindow(t *testing.T, interval, window time.Duration) {
|
||||
t.Helper()
|
||||
oldInterval, oldWindow := fetchRetryInterval, fetchRetryWindow
|
||||
fetchRetryInterval, fetchRetryWindow = interval, window
|
||||
t.Cleanup(func() { fetchRetryInterval, fetchRetryWindow = oldInterval, oldWindow })
|
||||
}
|
||||
|
||||
// A control-plane blip mid-fetch is survived: a 5xx on the first attempt is
|
||||
// retried and the second attempt's tarball extracts. This walks back the live
|
||||
// drill's failure, where the api pod rolled mid-fetch and the single attempt
|
||||
// died, failing the build Job.
|
||||
func TestFetchContextRetriesThroughBlip(t *testing.T) {
|
||||
shrinkFetchWindow(t, 10*time.Millisecond, time.Second)
|
||||
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
|
||||
var calls int32
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
if atomic.AddInt32(&calls, 1) == 1 {
|
||||
w.WriteHeader(http.StatusBadGateway) // the port is up, the API is not
|
||||
return
|
||||
}
|
||||
_, _ = w.Write(body)
|
||||
}))
|
||||
defer srv.Close()
|
||||
|
||||
dir := t.TempDir()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
var stderr bytes.Buffer
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr); code != 0 {
|
||||
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
|
||||
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
|
||||
}
|
||||
if !strings.Contains(stderr.String(), "retrying") {
|
||||
t.Fatalf("stderr %q does not mention the retry", stderr.String())
|
||||
}
|
||||
}
|
||||
|
||||
// The live drill's exact shape: the dial itself is refused (the api pod is
|
||||
// gone and no endpoint answers). A refused dial is retried like any other
|
||||
// transport failure, and once the face is back the fetch completes.
|
||||
func TestFetchContextRetriesRefusedDial(t *testing.T) {
|
||||
shrinkFetchWindow(t, 10*time.Millisecond, 5*time.Second)
|
||||
body := tgzBody(t, tarEntry{name: "Dockerfile", body: "FROM scratch\n"})
|
||||
|
||||
// Borrow a listen address, then close it: the first attempts dial into a
|
||||
// refused connection, exactly like a restarting control plane.
|
||||
probe := httptest.NewServer(http.HandlerFunc(func(http.ResponseWriter, *http.Request) {}))
|
||||
addr := strings.TrimPrefix(probe.URL, "http://")
|
||||
probe.Close()
|
||||
|
||||
dir := t.TempDir()
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
var stderr bytes.Buffer
|
||||
// Start the fetch; while the retry loop burns refused dials, bring the same
|
||||
// address back.
|
||||
result := make(chan int, 1)
|
||||
go func() {
|
||||
result <- cmdFetchContext([]string{"--url=http://" + addr + "/sub-1/context", "--out=" + dir}, io.Discard, &stderr)
|
||||
}()
|
||||
time.Sleep(100 * time.Millisecond) // let a handful of dials be refused
|
||||
ln, err := net.Listen("tcp", addr)
|
||||
if err != nil {
|
||||
t.Fatalf("rebind %s: %v", addr, err)
|
||||
}
|
||||
back := &http.Server{Handler: http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.Header.Get("Authorization") != "Bearer test-token" {
|
||||
w.WriteHeader(http.StatusUnauthorized)
|
||||
return
|
||||
}
|
||||
_, _ = w.Write(body)
|
||||
})}
|
||||
defer back.Close()
|
||||
go func() { _ = back.Serve(ln) }()
|
||||
|
||||
code := <-result
|
||||
if code != 0 {
|
||||
t.Fatalf("exit = %d, want 0 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
if got, err := os.ReadFile(filepath.Join(dir, "Dockerfile")); err != nil || string(got) != "FROM scratch\n" {
|
||||
t.Fatalf("extracted Dockerfile = (%q, %v)", got, err)
|
||||
}
|
||||
if !strings.Contains(stderr.String(), "retrying") {
|
||||
t.Fatalf("stderr %q does not mention the retry", stderr.String())
|
||||
}
|
||||
}
|
||||
|
||||
// A 4xx is an answer, not a blip: a missing/never-uploaded context fails
|
||||
// immediately — no retry loop burns the build's deadline on a terminal error.
|
||||
func TestFetchContextDoesNotRetry4xx(t *testing.T) {
|
||||
shrinkFetchWindow(t, 5*time.Millisecond, 200*time.Millisecond)
|
||||
var calls int32
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
atomic.AddInt32(&calls, 1)
|
||||
w.WriteHeader(http.StatusNotFound)
|
||||
}))
|
||||
defer srv.Close()
|
||||
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
var stderr bytes.Buffer
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
|
||||
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
if got := atomic.LoadInt32(&calls); got != 1 {
|
||||
t.Fatalf("server saw %d attempts, want exactly 1", got)
|
||||
}
|
||||
if strings.Contains(stderr.String(), "retrying") {
|
||||
t.Fatalf("stderr %q mentions a retry for a terminal 4xx", stderr.String())
|
||||
}
|
||||
}
|
||||
|
||||
// The retry is bounded: an internal face that stays down does not hang the
|
||||
// build pod; the window runs out and the fetch reports the exhausted retries.
|
||||
func TestFetchContextGivesUpAfterWindow(t *testing.T) {
|
||||
shrinkFetchWindow(t, 5*time.Millisecond, 60*time.Millisecond)
|
||||
var calls int32
|
||||
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
|
||||
atomic.AddInt32(&calls, 1)
|
||||
w.WriteHeader(http.StatusServiceUnavailable)
|
||||
}))
|
||||
defer srv.Close() // the face is up but never healthy: 503 forever
|
||||
|
||||
t.Setenv("FELIS_SERVICE_TOKEN", "test-token")
|
||||
var stderr bytes.Buffer
|
||||
start := time.Now()
|
||||
if code := cmdFetchContext([]string{"--url=" + srv.URL + "/sub-1/context", "--out=" + t.TempDir()}, io.Discard, &stderr); code != 1 {
|
||||
t.Fatalf("exit = %d, want 1 (stderr %q)", code, stderr.String())
|
||||
}
|
||||
if elapsed := time.Since(start); elapsed > 5*time.Second {
|
||||
t.Fatalf("gave up after %v; the window is supposed to bound it", elapsed)
|
||||
}
|
||||
if got := atomic.LoadInt32(&calls); got < 2 {
|
||||
t.Fatalf("server saw %d attempts, want at least one retry", got)
|
||||
}
|
||||
if !strings.Contains(stderr.String(), "retried for") {
|
||||
t.Fatalf("stderr %q does not report the exhausted retry window", stderr.String())
|
||||
}
|
||||
}
|
||||
+11
-16
@@ -20,18 +20,14 @@ const forwardingSecretEnv = "FELIS_FORWARDING_SECRET"
|
||||
// config/ and server.properties live under it.
|
||||
const defaultForwardingDataDir = "/data"
|
||||
|
||||
// fwd*Mode make the written config readable AND rewritable by the main server
|
||||
// container, whose UID we do not control (an arbitrary user image). The
|
||||
// initContainer runs as root (see buildStatefulSet) so it can write into a data
|
||||
// volume of unknown ownership; 0666/0777 then let a non-root Paper rewrite the
|
||||
// same files on boot.
|
||||
//
|
||||
// ponytail: relies on the initContainer running as root to write into a volume of
|
||||
// unknown ownership; that is how the operator schedules it. If that ever changes,
|
||||
// give the server pod an fsGroup so the shared volume is group-writable instead.
|
||||
// fwd*Mode are the modes the written config lands with. The initContainer runs as
|
||||
// the same uid as the server container (naming.GameUID, pinned by the operator in
|
||||
// the pod securityContext) after the prepare-data initContainer has handed the
|
||||
// whole volume to that uid, so owner read/write is all the server needs to rewrite
|
||||
// these files on boot and nothing else on the node gets write access to them.
|
||||
const (
|
||||
fwdFileMode os.FileMode = 0o666
|
||||
fwdDirMode os.FileMode = 0o777
|
||||
fwdFileMode os.FileMode = 0o644
|
||||
fwdDirMode os.FileMode = 0o755
|
||||
)
|
||||
|
||||
// cmdInitForwarding is the felis-image initContainer entrypoint that makes an
|
||||
@@ -96,9 +92,8 @@ func writePaperGlobal(dataDir, secret string) error {
|
||||
if err := os.MkdirAll(dir, fwdDirMode); err != nil {
|
||||
return fmt.Errorf("create %s: %w", dir, err)
|
||||
}
|
||||
// MkdirAll honours the process umask (root's is typically 022 → 0755); chmod
|
||||
// does not, and a non-root main container must be able to place/replace the
|
||||
// file in this directory on boot.
|
||||
// MkdirAll honours the process umask; chmod does not, so a directory an older
|
||||
// release left at 0777 is brought back to fwdDirMode here.
|
||||
if err := os.Chmod(dir, fwdDirMode); err != nil {
|
||||
return fmt.Errorf("chmod %s: %w", dir, err)
|
||||
}
|
||||
@@ -188,8 +183,8 @@ func upsertProperty(content []byte, key, value string) []byte {
|
||||
}
|
||||
|
||||
// writeFileMode writes data then forces the mode, since WriteFile honours the
|
||||
// umask (root's is typically 022 → 0644) but a non-root main container must be
|
||||
// able to rewrite these files on boot.
|
||||
// umask and leaves an existing file's mode alone: a file an older release wrote
|
||||
// world-writable (0666) is tightened back to fwdFileMode on the next boot.
|
||||
func writeFileMode(path string, data []byte) error {
|
||||
if err := os.WriteFile(path, data, fwdFileMode); err != nil {
|
||||
return fmt.Errorf("write %s: %w", path, err)
|
||||
|
||||
@@ -164,8 +164,9 @@ func TestUpsertPropertyAppends(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// The written files must be group/world writable so a non-root main container can
|
||||
// rewrite them. chmod semantics are POSIX-only, so this asserts on non-Windows.
|
||||
// The written files land at fwdFileMode: owner-writable for the game uid the init
|
||||
// shares with the server container, and no longer world-writable. chmod semantics
|
||||
// are POSIX-only, so this asserts on non-Windows.
|
||||
func TestWriteForwardingFileModes(t *testing.T) {
|
||||
if runtime.GOOS == "windows" {
|
||||
t.Skip("POSIX file modes not represented on Windows")
|
||||
|
||||
@@ -0,0 +1,117 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"io/fs"
|
||||
"os"
|
||||
"syscall"
|
||||
|
||||
"felis.lolicon.best/internal/naming"
|
||||
)
|
||||
|
||||
// cmdInitVolume is the felis-image `prepare-data` initContainer entrypoint: it
|
||||
// hands every entry of a server's world volume to the game uid/gid before the
|
||||
// server container starts. The operator runs the server itself as naming.GameUID,
|
||||
// so a world written by an earlier release (whose server ran as root), a restore
|
||||
// Job (which extracts as root), or a storage provisioner that creates the volume
|
||||
// root-owned would otherwise leave files the server cannot write — a world that
|
||||
// boots and then fails every save.
|
||||
//
|
||||
// fsGroup covers only part of this: kubelet applies it to volume types that
|
||||
// support ownership management, and a k3s local-path PV is a hostPath underneath,
|
||||
// which it skips. A walk from inside the pod works for every volume type.
|
||||
//
|
||||
// Only mismatched entries are touched, so a volume already owned by the game uid
|
||||
// costs one lstat per entry and no writes. The walk runs inside an os.Root at the
|
||||
// data dir and uses lchown, so a symlink a plugin planted is re-owned as a link
|
||||
// and never followed out of the volume.
|
||||
//
|
||||
// A single entry that cannot be chowned is reported and skipped: failing the pod
|
||||
// over one odd file would keep the whole server down, while the server itself
|
||||
// reports the one file it cannot write. Only an unreadable data dir fails.
|
||||
func cmdInitVolume(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("init-volume", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
dataDir := fs.String("data", defaultForwardingDataDir, "world volume mount to hand to the game uid")
|
||||
uid := fs.Int64("uid", naming.GameUID, "owner uid for every entry")
|
||||
gid := fs.Int64("gid", naming.GameGID, "owner gid for every entry")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
root, err := os.OpenRoot(*dataDir)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis init-volume: open %s: %v\n", *dataDir, err)
|
||||
return 1
|
||||
}
|
||||
defer root.Close()
|
||||
res, err := chownTree(root, int(*uid), int(*gid), root.Lchown)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis init-volume: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
for _, f := range res.failures {
|
||||
fmt.Fprintf(stderr, "felis init-volume: %s\n", f)
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis init-volume: %d entries checked, %d handed to %d:%d, %d failed\n",
|
||||
res.checked, res.changed, *uid, *gid, len(res.failures))
|
||||
return 0
|
||||
}
|
||||
|
||||
// chownResult tallies one walk; failures is capped so a volume of thousands of
|
||||
// unownable files cannot flood the pod log.
|
||||
type chownResult struct {
|
||||
checked int
|
||||
changed int
|
||||
failures []string
|
||||
}
|
||||
|
||||
const maxReportedChownFailures = 20
|
||||
|
||||
// chownTree walks root and calls chown on every entry (the root dir included)
|
||||
// whose owner is not uid:gid. It returns an error only when the root itself
|
||||
// cannot be read; per-entry failures are collected in the result.
|
||||
func chownTree(root *os.Root, uid, gid int, chown func(name string, uid, gid int) error) (chownResult, error) {
|
||||
var res chownResult
|
||||
fail := func(name string, err error) {
|
||||
if len(res.failures) < maxReportedChownFailures {
|
||||
res.failures = append(res.failures, fmt.Sprintf("%s: %v", name, err))
|
||||
} else if len(res.failures) == maxReportedChownFailures {
|
||||
res.failures = append(res.failures, "further failures not listed")
|
||||
}
|
||||
}
|
||||
err := fs.WalkDir(root.FS(), ".", func(name string, d fs.DirEntry, walkErr error) error {
|
||||
if walkErr != nil {
|
||||
if name == "." {
|
||||
return walkErr
|
||||
}
|
||||
fail(name, walkErr)
|
||||
// A directory that cannot be listed is skipped as a whole; a file
|
||||
// error has nothing below it to skip.
|
||||
if d != nil && d.IsDir() {
|
||||
return fs.SkipDir
|
||||
}
|
||||
return nil
|
||||
}
|
||||
res.checked++
|
||||
info, err := d.Info()
|
||||
if err != nil {
|
||||
fail(name, err)
|
||||
return nil
|
||||
}
|
||||
if st, ok := info.Sys().(*syscall.Stat_t); ok && int(st.Uid) == uid && int(st.Gid) == gid {
|
||||
return nil
|
||||
}
|
||||
if err := chown(name, uid, gid); err != nil {
|
||||
if !errors.Is(err, fs.ErrNotExist) { // gone mid-walk: nothing left to own
|
||||
fail(name, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
res.changed++
|
||||
return nil
|
||||
})
|
||||
return res, err
|
||||
}
|
||||
@@ -0,0 +1,124 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
"slices"
|
||||
"strconv"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// openTree builds a small world under a temp dir: nested dirs, a file, and a
|
||||
// symlink pointing out of the volume that the walk must not follow.
|
||||
func openTree(t *testing.T) *os.Root {
|
||||
t.Helper()
|
||||
dir := t.TempDir()
|
||||
for _, d := range []string{"world/region", "plugins"} {
|
||||
if err := os.MkdirAll(filepath.Join(dir, d), 0o755); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
for _, f := range []string{"level.dat", "world/region/r.0.0.mca"} {
|
||||
if err := os.WriteFile(filepath.Join(dir, f), []byte("x"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
outside := t.TempDir()
|
||||
if err := os.Symlink(outside, filepath.Join(dir, "plugins", "escape")); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
root, err := os.OpenRoot(dir)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Cleanup(func() { root.Close() })
|
||||
return root
|
||||
}
|
||||
|
||||
// Every entry owned by someone else is handed over, the root dir included, and a
|
||||
// symlink is re-owned as a link rather than walked into.
|
||||
func TestChownTreeHandsOverMismatchedEntries(t *testing.T) {
|
||||
if runtime.GOOS == "windows" {
|
||||
t.Skip("POSIX ownership not represented on Windows")
|
||||
}
|
||||
root := openTree(t)
|
||||
var got []string
|
||||
res, err := chownTree(root, os.Getuid()+1, os.Getgid(), func(name string, uid, gid int) error {
|
||||
got = append(got, name)
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("chownTree: %v", err)
|
||||
}
|
||||
want := []string{".", "level.dat", "plugins", "plugins/escape", "world", "world/region", "world/region/r.0.0.mca"}
|
||||
slices.Sort(got)
|
||||
if !slices.Equal(got, want) {
|
||||
t.Errorf("chowned %v, want %v", got, want)
|
||||
}
|
||||
if res.changed != len(want) || res.checked != len(want) || len(res.failures) != 0 {
|
||||
t.Errorf("result = %+v, want %d checked and changed, no failures", res, len(want))
|
||||
}
|
||||
}
|
||||
|
||||
// A volume already owned by the game uid costs no chown at all: this is the steady
|
||||
// state every restart after the first one hits.
|
||||
func TestChownTreeSkipsMatchingOwner(t *testing.T) {
|
||||
if runtime.GOOS == "windows" {
|
||||
t.Skip("POSIX ownership not represented on Windows")
|
||||
}
|
||||
root := openTree(t)
|
||||
calls := 0
|
||||
res, err := chownTree(root, os.Getuid(), os.Getgid(), func(string, int, int) error {
|
||||
calls++
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("chownTree: %v", err)
|
||||
}
|
||||
if calls != 0 || res.changed != 0 {
|
||||
t.Errorf("chown called %d times on an already-owned tree (result %+v)", calls, res)
|
||||
}
|
||||
}
|
||||
|
||||
// One entry that refuses the chown is reported and the walk carries on: a single
|
||||
// odd file must not keep the whole server from starting.
|
||||
func TestChownTreeContinuesPastFailures(t *testing.T) {
|
||||
if runtime.GOOS == "windows" {
|
||||
t.Skip("POSIX ownership not represented on Windows")
|
||||
}
|
||||
root := openTree(t)
|
||||
res, err := chownTree(root, os.Getuid()+1, os.Getgid(), func(name string, uid, gid int) error {
|
||||
if name == "level.dat" {
|
||||
return os.ErrPermission
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("chownTree: %v", err)
|
||||
}
|
||||
if len(res.failures) != 1 || res.changed != 6 {
|
||||
t.Errorf("result = %+v, want 1 failure and 6 changed", res)
|
||||
}
|
||||
}
|
||||
|
||||
// Against a real directory the owner already matches, so the command succeeds
|
||||
// without needing CAP_CHOWN — the path every test runner (non-root) can take.
|
||||
func TestCmdInitVolumeOwnedTree(t *testing.T) {
|
||||
if runtime.GOOS == "windows" {
|
||||
t.Skip("POSIX ownership not represented on Windows")
|
||||
}
|
||||
dir := t.TempDir()
|
||||
if err := os.WriteFile(filepath.Join(dir, "server.properties"), []byte("x"), 0o644); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var out, errb bytes.Buffer
|
||||
code := cmdInitVolume([]string{"--data", dir, "--uid", strconv.Itoa(os.Getuid()), "--gid", strconv.Itoa(os.Getgid())}, &out, &errb)
|
||||
if code != 0 {
|
||||
t.Fatalf("exit %d, stderr %q", code, errb.String())
|
||||
}
|
||||
if code := cmdInitVolume([]string{"--data", filepath.Join(dir, "missing")}, &out, &errb); code != 1 {
|
||||
t.Errorf("missing data dir exit = %d, want 1", code)
|
||||
}
|
||||
}
|
||||
+70
-28
@@ -26,8 +26,10 @@ func (m *multiFlag) Set(v string) error {
|
||||
// felis-reaper identity only when the retention reaper is enabled, gated with
|
||||
// its CronJob), the weak build/restore Job SAs, the build/minecraft
|
||||
// NetworkPolicies, and the running control-plane workloads (felis-api/operator
|
||||
// Deployments + the in-cluster registry Deployment/Service/PVC) — as a single
|
||||
// multi-document YAML stream on stdout, ready for `kubectl apply -f -`.
|
||||
// Deployments + the in-cluster registry Deployment/Service/PVC + the
|
||||
// world-archive PVC that backs backup/restore, unless --backup-pvc is emptied)
|
||||
// — as a single multi-document YAML stream on stdout, ready for
|
||||
// `kubectl apply -f -`.
|
||||
//
|
||||
// It is a pure renderer: it never contacts a cluster and holds no credentials.
|
||||
// --velocity-cidr records the proxy host addresses allowed by the game NetworkPolicy.
|
||||
@@ -43,14 +45,22 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
|
||||
registryPort := fs.Int("registry-port", 5000, "port the in-cluster registry listens on")
|
||||
panelNodePort := fs.Int("panel-node-port", int(platform.DefaultPanelNodePort), "NodePort that exposes the built-in HTTPS panel/API origin")
|
||||
felisImage := fs.String("felis-image", "", "container image the felis-api/operator Deployments run, also passed through as FELIS_IMAGE (REQUIRED)")
|
||||
registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry:2)")
|
||||
backupPVC := fs.String("backup-pvc", "", "name of the backup PVC advertised to the restore executor via FELIS_BACKUP_PVC (default none = restore endpoint returns 503)")
|
||||
worldsHostPath := fs.String("worlds-host-path", "", "node directory under which each world PVC is visible as <path>/<pvc>; enables the reaper CronJob (requires --backup-pvc and --archive-local-path)")
|
||||
registryImage := fs.String("registry-image", "", "in-cluster registry image (default: registry 2.8.3, pinned by digest)")
|
||||
backupPVC := fs.String("backup-pvc", "felis-backups", "name of the world-archive PVC this bundle renders in the Minecraft namespace and advertises to the backup/restore executors via FELIS_BACKUP_PVC (default: felis-backups; pass an empty value to render none, leaving backup/restore answering 503)")
|
||||
worldsHostPath := fs.String("worlds-host-path", "", "node directory the reaper reads worlds from: each world PVC resolves as <path>/<pvc>, or as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (k3s storage root: /var/lib/rancher/k3s/storage); enables the reaper CronJob (requires --archive-local-path and a non-empty --backup-pvc)")
|
||||
archiveLocalPath := fs.String("archive-local-path", "", "path the backup PVC is mounted at in the reaper CronJob; MUST equal felis.toml [archive] local_path")
|
||||
registryStorage := fs.String("registry-storage", "", "capacity the registry PVC requests (default 10Gi; k3s local-path does not enforce it)")
|
||||
uploadsStorage := fs.String("uploads-storage", "", "capacity the uploads PVC requests (default 5Gi; k3s local-path does not enforce it)")
|
||||
backupStorage := fs.String("backup-storage", "", "capacity the world-archive PVC requests (default 10Gi; k3s local-path does not enforce it)")
|
||||
reaperNode := fs.String("reaper-node", "", "node that holds --worlds-host-path: pins the reaper CronJob's pod there via nodeSelector kubernetes.io/hostname (multi-node clusters need this, or the reaper may schedule where the hostPath is empty)")
|
||||
var velocityCIDRs multiFlag
|
||||
fs.Var(&velocityCIDRs, "velocity-cidr", "CIDR of a Velocity proxy host allowed to reach game port 25565 (repeatable, REQUIRED)")
|
||||
var packageCIDRs multiFlag
|
||||
fs.Var(&packageCIDRs, "package-cidr", "CIDR of a package mirror build Pods may reach (repeatable; default none = no internet egress)")
|
||||
var serverDenyCIDRs multiFlag
|
||||
fs.Var(&serverDenyCIDRs, "server-egress-deny-cidr", "extra CIDR game server pods may never reach, e.g. the node's public address (repeatable)")
|
||||
var serverAllowCIDRs multiFlag
|
||||
fs.Var(&serverAllowCIDRs, "server-egress-allow-cidr", "private CIDR game server pods may reach despite the private-range block, e.g. a LAN database (repeatable)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
@@ -71,7 +81,9 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
|
||||
"(the felis-api/operator Deployments run it and it is passed through as FELIS_IMAGE, e.g. --felis-image registry.felis.svc:5000/felis:v1)")
|
||||
return 2
|
||||
}
|
||||
for _, cidr := range append(append([]string{}, velocityCIDRs...), packageCIDRs...) {
|
||||
allCIDRs := append(append([]string{}, velocityCIDRs...), packageCIDRs...)
|
||||
allCIDRs = append(append(allCIDRs, serverDenyCIDRs...), serverAllowCIDRs...)
|
||||
for _, cidr := range allCIDRs {
|
||||
if _, _, err := net.ParseCIDR(cidr); err != nil {
|
||||
fmt.Fprintf(stderr, "felis manifests: invalid CIDR %q: %v\n", cidr, err)
|
||||
return 2
|
||||
@@ -81,39 +93,56 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
|
||||
fmt.Fprintf(stderr, "felis manifests: --panel-node-port must be in Kubernetes NodePort range 30000-32767 (got %d)\n", *panelNodePort)
|
||||
return 2
|
||||
}
|
||||
// The node pin exists only for the reaper's hostPath: naming a node without the
|
||||
// worlds root would be silently dropped (no CronJob renders), so fail loud like
|
||||
// the storage-trio check below.
|
||||
if *reaperNode != "" && *worldsHostPath == "" {
|
||||
fmt.Fprintln(stderr, "felis manifests: --reaper-node requires --worlds-host-path "+
|
||||
"(it pins the reaper CronJob, which renders only with the retention storage trio)")
|
||||
return 2
|
||||
}
|
||||
|
||||
// Retention/reaper rendering is opt-in and needs all three storage coordinates
|
||||
// together: where worlds live (to read+archive them), the backup PVC (to write
|
||||
// archives into), and the path it is mounted at (which MUST equal felis.toml
|
||||
// [archive] local_path so tarLocal's absolute archive refs resolve). A partial
|
||||
// configuration is almost certainly an operator mistake, so fail loud rather than
|
||||
// silently drop retention. Asking for it without the other two is rejected; an
|
||||
// empty trio renders the bundle WITHOUT the reaper and says so.
|
||||
// Retention/reaper rendering is opt-in and needs a storage topology together:
|
||||
// where worlds live (to read+archive them), a backup PVC (to write archives
|
||||
// into — rendered from --backup-pvc), and the path it is mounted at (which MUST
|
||||
// equal felis.toml [archive] local_path so tarLocal's absolute archive refs
|
||||
// resolve). A partial configuration is almost certainly an operator mistake, so
|
||||
// fail loud rather than silently drop retention or render a reaper with nowhere
|
||||
// to write. The backup PVC itself defaults to felis-backups (it is what makes a
|
||||
// default install's backup endpoint work at all); retention additionally needs
|
||||
// --worlds-host-path.
|
||||
if *worldsHostPath != "" {
|
||||
if *backupPVC == "" || *archiveLocalPath == "" {
|
||||
fmt.Fprintln(stderr, "felis manifests: --worlds-host-path enables the reaper CronJob and requires "+
|
||||
"--backup-pvc and --archive-local-path too (--archive-local-path must equal felis.toml [archive] local_path)")
|
||||
"--archive-local-path (must equal felis.toml [archive] local_path) and a non-empty --backup-pvc "+
|
||||
"(the archive store; default felis-backups)")
|
||||
return 2
|
||||
}
|
||||
// The reaper WILL render. Two deployment preconditions this generator cannot
|
||||
// check would SILENTLY turn retention into a no-op if unmet — surface them as
|
||||
// The reaper WILL render. Two deployment facts this generator cannot check
|
||||
// would silently turn retention into a no-op if unmet — surface them as
|
||||
// loudly as the fail-closed cases above, so an operator is never left with a
|
||||
// reaper that reaps an empty directory. (Both are also in the WorldsHostPath
|
||||
// flag/field docs, but nobody deploying from stdout reads those.)
|
||||
// reaper that reaps nothing. (Both are also in the WorldsHostPath flag/field
|
||||
// docs, but nobody deploying from stdout reads those.)
|
||||
pin := "the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but on a " +
|
||||
"multi-node cluster you MUST pass --reaper-node <name> (or add a nodeSelector) for the node holding the " +
|
||||
"worlds, or the reaper may schedule where the hostPath is empty"
|
||||
if *reaperNode != "" {
|
||||
pin = fmt.Sprintf("the CronJob and its worlds-root PV are pinned to node %q via kubernetes.io/hostname — "+
|
||||
"keep this pointed at the node that actually holds the world volumes", *reaperNode)
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis manifests: note: rendering the retention reaper CronJob (worlds hostPath %q). "+
|
||||
"Two preconditions are NOT verified here:\n"+
|
||||
" - each world PVC must be visible at %s/<pvc> on the node: a stock local-path-provisioner lays "+
|
||||
"volumes under PV-name paths (.../pvc-<uuid>_<ns>_<pvc>/), so unless the worlds StorageClass is "+
|
||||
"arranged to expose <path>/<pvc>, the reaper tars an empty directory;\n"+
|
||||
" - the CronJob sets NO nodeSelector: a single-node starter pins it to the worlds implicitly, but "+
|
||||
"on a multi-node cluster you MUST add a nodeSelector for the node holding the worlds, or the reaper "+
|
||||
"may schedule where the hostPath is empty.\n", *worldsHostPath, *worldsHostPath)
|
||||
"These points are NOT verified here:\n"+
|
||||
" - the node's world volumes must actually live below %s: the reaper resolves a world as "+
|
||||
"%s/<pvc>, then as the stock local-path directory <path>/<pv-name>_<ns>_<pvc-name> (what k3s "+
|
||||
"writes under /var/lib/rancher/k3s/storage). Any other provisioner needs its volumes exposed as "+
|
||||
"<path>/<pvc>, or each candidate's archive fails and the world is preserved;\n"+
|
||||
" - %s.\n", *worldsHostPath, *worldsHostPath, *worldsHostPath, pin)
|
||||
} else {
|
||||
fmt.Fprintln(stderr, "felis manifests: note: retention reaper CronJob not rendered "+
|
||||
"(pass --worlds-host-path, --backup-pvc and --archive-local-path to enable it)")
|
||||
"(pass --worlds-host-path and --archive-local-path — the archive PVC defaults to felis-backups — to enable it)")
|
||||
}
|
||||
|
||||
out, err := platform.RenderYAML(platform.Params{
|
||||
params := platform.Params{
|
||||
ControlNamespace: *controlNS,
|
||||
MinecraftNamespace: *minecraftNS,
|
||||
BuildNamespace: *buildNS,
|
||||
@@ -124,10 +153,23 @@ func cmdManifests(args []string, stdout, stderr io.Writer) int {
|
||||
RegistryImage: *registryImage,
|
||||
BackupPVC: *backupPVC,
|
||||
WorldsHostPath: *worldsHostPath,
|
||||
ReaperNode: *reaperNode,
|
||||
ArchiveLocalPath: *archiveLocalPath,
|
||||
VelocityCIDRs: []string(velocityCIDRs),
|
||||
PackageSourceCIDRs: []string(packageCIDRs),
|
||||
})
|
||||
|
||||
ServerEgressDenyCIDRs: []string(serverDenyCIDRs),
|
||||
ServerEgressAllowCIDRs: []string(serverAllowCIDRs),
|
||||
|
||||
RegistryStorage: *registryStorage,
|
||||
UploadsStorage: *uploadsStorage,
|
||||
BackupStorage: *backupStorage,
|
||||
}
|
||||
if err := params.Validate(); err != nil {
|
||||
fmt.Fprintf(stderr, "felis manifests: %v\n", err)
|
||||
return 2
|
||||
}
|
||||
out, err := platform.RenderYAML(params)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis manifests: render: %v\n", err)
|
||||
return 1
|
||||
|
||||
@@ -77,6 +77,10 @@ func TestManifestsRendersBundle(t *testing.T) {
|
||||
"10.0.0.5/32",
|
||||
// The felis image flows through to the Deployments.
|
||||
"registry.felis.svc:5000/felis:v1",
|
||||
// Backup works out of the box: the archive PVC renders and the api gets
|
||||
// the env that wires the backup/restore executors to it.
|
||||
"name: felis-backups",
|
||||
"name: FELIS_BACKUP_PVC",
|
||||
} {
|
||||
if !strings.Contains(text, want) {
|
||||
t.Errorf("rendered bundle missing %q", want)
|
||||
@@ -95,15 +99,19 @@ func TestManifestsRendersBundle(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestManifestsReaperRequiresTrio proves --worlds-host-path is a fail-loud opt-in:
|
||||
// asking for the reaper without the backup PVC and its mount path (which must equal
|
||||
// [archive] local_path) is rejected rather than silently dropping retention.
|
||||
func TestManifestsReaperRequiresTrio(t *testing.T) {
|
||||
// TestManifestsReaperRequiresStorage proves --worlds-host-path is a fail-loud
|
||||
// opt-in: asking for the reaper without a writable archive store (the backup PVC,
|
||||
// which defaults to felis-backups but can be emptied) and its mount path (which
|
||||
// must equal [archive] local_path) is rejected rather than silently dropping
|
||||
// retention or deleting worlds it could not archive first.
|
||||
func TestManifestsReaperRequiresStorage(t *testing.T) {
|
||||
base := []string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32", "--worlds-host-path", "/var/lib/felis/worlds"}
|
||||
for _, extra := range [][]string{
|
||||
{}, // neither backup-pvc nor archive-local-path
|
||||
{"--backup-pvc", "felis-backups"}, // missing archive-local-path
|
||||
{"--archive-local-path", "/backups"}, // missing backup-pvc
|
||||
{}, // missing archive-local-path (backup-pvc defaults)
|
||||
{"--backup-pvc", "other"}, // still missing archive-local-path
|
||||
// A reaper with no archive store would have nowhere to write the archive
|
||||
// it must verify before deleting a world; emptying the PVC is rejected.
|
||||
{"--archive-local-path", "/backups", "--backup-pvc="},
|
||||
} {
|
||||
var out, errBuf bytes.Buffer
|
||||
code := run(append(append([]string{}, base...), extra...), &out, &errBuf)
|
||||
@@ -119,6 +127,23 @@ func TestManifestsReaperRequiresTrio(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestManifestsBackupPVCOptOut proves --backup-pvc= renders a bundle with no
|
||||
// archive store at all: no PVC and no FELIS_BACKUP_PVC env, so backup/restore
|
||||
// answer 503 instead of pointing Jobs at a claim nobody provisions.
|
||||
func TestManifestsBackupPVCOptOut(t *testing.T) {
|
||||
var out, errBuf bytes.Buffer
|
||||
code := run([]string{"manifests", "--felis-image", "reg/felis:test",
|
||||
"--velocity-cidr", "10.0.0.5/32", "--backup-pvc="}, &out, &errBuf)
|
||||
if code != 0 {
|
||||
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
|
||||
}
|
||||
for _, absent := range []string{"felis-backups", "FELIS_BACKUP_PVC"} {
|
||||
if strings.Contains(out.String(), absent) {
|
||||
t.Errorf("--backup-pvc= bundle must not contain %q", absent)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestManifestsRendersReaper proves the happy path with the full retention trio:
|
||||
// a batch/v1 CronJob is emitted, named felis-reaper, mounting the backup PVC at the
|
||||
// supplied archive path.
|
||||
@@ -150,9 +175,71 @@ func TestManifestsRendersReaper(t *testing.T) {
|
||||
// this generator cannot verify (else a misarranged hostPath silently no-ops
|
||||
// retention): the <path>/<pvc> arrangement-dependency and the multi-node
|
||||
// nodeSelector hazard.
|
||||
for _, want := range []string{"local-path-provisioner", "nodeSelector"} {
|
||||
for _, want := range []string{"local-path", "nodeSelector"} {
|
||||
if !strings.Contains(errBuf.String(), want) {
|
||||
t.Errorf("reaper render must warn operators about %q on stderr, got %q", want, errBuf.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestManifestsReaperNodePin: --reaper-node pins the rendered CronJob's pod via
|
||||
// kubernetes.io/hostname and replaces the "no nodeSelector" hazard note with the
|
||||
// pin confirmation; using it without the worlds root is a fail-loud 2.
|
||||
func TestManifestsReaperNodePin(t *testing.T) {
|
||||
var out, errBuf bytes.Buffer
|
||||
code := run([]string{
|
||||
"manifests",
|
||||
"--felis-image", "registry.felis.svc:5000/felis:v1",
|
||||
"--velocity-cidr", "10.0.0.5/32",
|
||||
"--worlds-host-path", "/var/lib/felis/worlds",
|
||||
"--archive-local-path", "/backups",
|
||||
"--reaper-node", "node-a",
|
||||
}, &out, &errBuf)
|
||||
if code != 0 {
|
||||
t.Fatalf("exit code = %d, want 0; stderr=%q", code, errBuf.String())
|
||||
}
|
||||
for _, want := range []string{
|
||||
"kubernetes.io/hostname: node-a",
|
||||
} {
|
||||
if !strings.Contains(out.String(), want) {
|
||||
t.Errorf("pinned render missing %q", want)
|
||||
}
|
||||
}
|
||||
if !strings.Contains(errBuf.String(), "node-a") {
|
||||
t.Errorf("stderr must confirm the pin, got %q", errBuf.String())
|
||||
}
|
||||
|
||||
var out2, err2 bytes.Buffer
|
||||
if code := run([]string{
|
||||
"manifests",
|
||||
"--felis-image", "registry.felis.svc:5000/felis:v1",
|
||||
"--velocity-cidr", "10.0.0.5/32",
|
||||
"--reaper-node", "node-a",
|
||||
}, &out2, &err2); code != 2 {
|
||||
t.Errorf("--reaper-node without --worlds-host-path: exit = %d, want 2", code)
|
||||
}
|
||||
}
|
||||
|
||||
// TestManifestsStorageSizes proves the PVC size flags reach the rendered claims
|
||||
// and a size the API server would reject fails before anything is applied.
|
||||
func TestManifestsStorageSizes(t *testing.T) {
|
||||
var out, errBuf bytes.Buffer
|
||||
code := run([]string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32",
|
||||
"--registry-storage", "40Gi", "--uploads-storage", "8Gi"}, &out, &errBuf)
|
||||
if code != 0 {
|
||||
t.Fatalf("exit code = %d, stderr = %s", code, errBuf.String())
|
||||
}
|
||||
for _, want := range []string{"storage: 40Gi", "storage: 8Gi"} {
|
||||
if !strings.Contains(out.String(), want) {
|
||||
t.Errorf("bundle lacks %q", want)
|
||||
}
|
||||
}
|
||||
|
||||
out.Reset()
|
||||
errBuf.Reset()
|
||||
code = run([]string{"manifests", "--felis-image", "reg/felis:test", "--velocity-cidr", "10.0.0.5/32",
|
||||
"--registry-storage", "lots"}, &out, &errBuf)
|
||||
if code == 0 || !strings.Contains(errBuf.String(), "registry storage") {
|
||||
t.Errorf("--registry-storage lots: exit %d, stderr %q; want a refusal naming the flag", code, errBuf.String())
|
||||
}
|
||||
}
|
||||
@@ -7,15 +7,24 @@ import (
|
||||
"io"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/dbbackup"
|
||||
"felis.lolicon.best/internal/store"
|
||||
)
|
||||
|
||||
// cmdMigrate implements `felis migrate up`: load config, open the database, and
|
||||
// apply every pending embedded migration under the advisory lock (spec §6).
|
||||
//
|
||||
// Migrations only roll forward, and some drop data (0017_drop_password), so a
|
||||
// database that already holds a schema and has migrations pending is bundled
|
||||
// first (internal/dbbackup, label pre-migrate). A failed snapshot stops the
|
||||
// upgrade; -no-backup is the explicit way past it, e.g. for an external
|
||||
// database whose server is newer than the host's pg_dump.
|
||||
func cmdMigrate(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("migrate", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
backupDir := fs.String("backup-dir", dbbackup.DefaultDir, "where the pre-migration snapshot goes")
|
||||
noBackup := fs.Bool("no-backup", false, "apply pending migrations without snapshotting the database first")
|
||||
// The "up" verb precedes any flags (felis migrate up -config path). Go's
|
||||
// flag.Parse stops at the first non-flag token and would never see a flag
|
||||
// placed after "up", silently falling back to the default -config. Pull the
|
||||
@@ -48,6 +57,18 @@ func cmdMigrate(args []string, stdout, stderr io.Writer) int {
|
||||
return 1
|
||||
}
|
||||
|
||||
if !*noBackup {
|
||||
path, err := preMigrateBackup(ctx, drv, migrations, cfg.Database.URL, *backupDir, stderr)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis migrate: pre-migration backup failed, nothing applied: %v\n", err)
|
||||
fmt.Fprintln(stderr, " fix the backup, or re-run with -no-backup to migrate without one")
|
||||
return 1
|
||||
}
|
||||
if path != "" {
|
||||
fmt.Fprintf(stdout, "felis migrate: database snapshot %s\n", path)
|
||||
}
|
||||
}
|
||||
|
||||
applied, err := store.Up(ctx, drv, migrations)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis migrate: %v\n", err)
|
||||
@@ -60,3 +81,33 @@ func cmdMigrate(args []string, stdout, stderr io.Writer) int {
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// preMigrateBackup bundles the database when it already carries a schema and
|
||||
// some of migrations are not applied yet, and returns the bundle's path ("" when
|
||||
// there was nothing to protect: a fresh database, or nothing pending).
|
||||
func preMigrateBackup(ctx context.Context, drv store.Driver, migrations []store.Migration, dbURL, dir string, log io.Writer) (string, error) {
|
||||
if err := drv.EnsureVersionTable(ctx); err != nil {
|
||||
return "", fmt.Errorf("ensure version table: %w", err)
|
||||
}
|
||||
done, err := drv.AppliedVersions(ctx)
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("read applied versions: %w", err)
|
||||
}
|
||||
if len(done) == 0 || !hasPending(done, migrations) {
|
||||
return "", nil
|
||||
}
|
||||
return dbbackup.Backup(ctx, dbbackup.BackupOptions{
|
||||
DatabaseURL: dbURL, Dir: dir, Label: dbbackup.LabelPreMigrate,
|
||||
Keep: defaultKeep[dbbackup.LabelPreMigrate], StateDir: dbbackup.DefaultStateDir,
|
||||
Version: resolvedVersion(), Log: log, Record: true,
|
||||
})
|
||||
}
|
||||
|
||||
func hasPending(done map[int]struct{}, migrations []store.Migration) bool {
|
||||
for _, m := range migrations {
|
||||
if _, ok := done[m.Version]; !ok {
|
||||
return true
|
||||
}
|
||||
}
|
||||
return false
|
||||
}
|
||||
@@ -0,0 +1,125 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/signal"
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/build"
|
||||
"felis.lolicon.best/internal/imagepush"
|
||||
"felis.lolicon.best/internal/registrygate"
|
||||
)
|
||||
|
||||
// defaultBuildToolsStatus is where mirror-build-tools records its last run; the
|
||||
// watchdog reads it to tell a vulnerability DB that stopped refreshing.
|
||||
const defaultBuildToolsStatus = "/var/lib/felis/build-tools/status.json"
|
||||
|
||||
// cmdMirrorBuildTools copies the build lane's tools (build.Tools: the kaniko and
|
||||
// trivy images, Trivy's vulnerability and Java DBs) from upstream into the
|
||||
// platform registry, where build Jobs pull them. deploy/bootstrap.sh runs it at
|
||||
// install and from felis-build-tools.timer twice a day, which is what keeps the
|
||||
// DBs fresh; a root shell can run it the same way to refresh now.
|
||||
//
|
||||
// It writes as the platform principal through the node's loopback hostPort, the
|
||||
// same way the installer pushes, reading the token from the environment or from
|
||||
// /etc/felis/secrets.env.
|
||||
func cmdMirrorBuildTools(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("mirror-build-tools", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
endpoint := fs.String("endpoint", "127.0.0.1:5000", "host[:port] of the registry to write to (plain HTTP)")
|
||||
only := fs.String("only", "", "comma-separated tool names to copy (default: all of "+toolNames()+")")
|
||||
status := fs.String("status", defaultBuildToolsStatus, `file to record the run in ("" records nothing)`)
|
||||
secrets := fs.String("secrets-env", "/etc/felis/secrets.env", "installer secrets file holding REGISTRY_PLATFORM_TOKEN, read when FELIS_REGISTRY_PASSWORD is unset")
|
||||
platformFlag := fs.String("platform", "", "os/arch of the images to copy (default: this machine's)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return 0
|
||||
}
|
||||
return 2
|
||||
}
|
||||
tools, err := selectTools(*only)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis mirror-build-tools: %v\n", err)
|
||||
return 2
|
||||
}
|
||||
if err := loadEnvFile(*secrets); err != nil {
|
||||
fmt.Fprintf(stderr, "felis mirror-build-tools: read %s: %v\n", *secrets, err)
|
||||
return 1
|
||||
}
|
||||
user, pass := os.Getenv("FELIS_REGISTRY_USERNAME"), os.Getenv("FELIS_REGISTRY_PASSWORD")
|
||||
if pass == "" {
|
||||
user, pass = registrygate.PrincipalPlatform, os.Getenv("REGISTRY_PLATFORM_TOKEN")
|
||||
}
|
||||
if pass == "" {
|
||||
fmt.Fprintln(stderr, "felis mirror-build-tools: no registry credential: set FELIS_REGISTRY_PASSWORD or run as root on the node (REGISTRY_PLATFORM_TOKEN in /etc/felis/secrets.env)")
|
||||
return 2
|
||||
}
|
||||
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
p := &imagepush.Pusher{Scheme: "http", Username: user, Password: pass, Log: stdout}
|
||||
src := &imagepush.Source{Platform: *platformFlag}
|
||||
started := time.Now()
|
||||
var failed []string
|
||||
for _, t := range tools {
|
||||
dst := strings.TrimSuffix(*endpoint, "/") + "/" + t.Mirror
|
||||
if _, err := p.Mirror(ctx, src, t.Source, dst); err != nil {
|
||||
fmt.Fprintf(stderr, "felis mirror-build-tools: %s: %v\n", t.Name, err)
|
||||
failed = append(failed, t.Name+": "+err.Error())
|
||||
}
|
||||
}
|
||||
if *status != "" {
|
||||
st, err := imagepush.ReadMirrorStatus(*status)
|
||||
if err != nil || st == nil {
|
||||
st = &imagepush.MirrorStatus{}
|
||||
}
|
||||
st.LastAttempt = started
|
||||
st.LastError = strings.Join(failed, "; ")
|
||||
if len(failed) == 0 {
|
||||
st.LastSuccess = started
|
||||
}
|
||||
if err := imagepush.WriteMirrorStatus(*status, *st); err != nil {
|
||||
fmt.Fprintf(stderr, "felis mirror-build-tools: record %s: %v\n", *status, err)
|
||||
}
|
||||
}
|
||||
if len(failed) > 0 {
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func toolNames() string {
|
||||
var names []string
|
||||
for _, t := range build.Tools {
|
||||
names = append(names, t.Name)
|
||||
}
|
||||
return strings.Join(names, ",")
|
||||
}
|
||||
|
||||
func selectTools(only string) ([]build.Tool, error) {
|
||||
if only == "" {
|
||||
return build.Tools, nil
|
||||
}
|
||||
var out []build.Tool
|
||||
for _, name := range strings.Split(only, ",") {
|
||||
name = strings.TrimSpace(name)
|
||||
found := false
|
||||
for _, t := range build.Tools {
|
||||
if t.Name == name {
|
||||
out = append(out, t)
|
||||
found = true
|
||||
}
|
||||
}
|
||||
if !found {
|
||||
return nil, fmt.Errorf("unknown tool %q (known: %s)", name, toolNames())
|
||||
}
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
+71
-17
@@ -29,7 +29,12 @@ import (
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"net"
|
||||
"net/http"
|
||||
"os"
|
||||
"os/signal"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/api"
|
||||
"felis.lolicon.best/internal/config"
|
||||
@@ -37,21 +42,27 @@ import (
|
||||
|
||||
// nanoStubRepo satisfies api.Repo but implements only the one method handleHasJoined calls.
|
||||
// The reclaim username blacklist is a felis-api/DB concern; a nano host has no Postgres, so
|
||||
// nothing is barred here. ponytail: a real blacklist would need the very DB nano exists to
|
||||
// nothing is barred here. A real blacklist would need the very DB nano exists to
|
||||
// avoid — YAGNI until a nano host grows a reclaim store.
|
||||
type nanoStubRepo struct{ api.Repo }
|
||||
|
||||
func (nanoStubRepo) IsUsernameBlacklisted(context.Context, string) (bool, error) { return false, nil }
|
||||
|
||||
// nanoDefaultListen is loopback because hasJoined carries no auth token (Velocity speaks the
|
||||
// vanilla sessionserver protocol), so a public bind is an open auth relay: anyone can point
|
||||
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity reaches
|
||||
// 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
|
||||
const nanoDefaultListen = "127.0.0.1:8081"
|
||||
|
||||
// nanoLogURIMax is room for a real hasJoined query (a 16-character name, a 41-character
|
||||
// serverId, an address) several times over.
|
||||
const nanoLogURIMax = 256
|
||||
|
||||
func cmdNano(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("nano", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (reads [[auth_source]])")
|
||||
// Loopback default: hasJoined carries no auth token (authlib speaks the vanilla
|
||||
// sessionserver protocol), so a public bind is an open auth relay — anyone can point
|
||||
// their proxy at it and spend this host's egress IP on Mojang. A same-host Velocity
|
||||
// reaches 127.0.0.1; serving an off-host proxy is an explicit -listen opt-in.
|
||||
listen := fs.String("listen", "127.0.0.1:8081", "listen address for the hasJoined endpoint")
|
||||
listen := fs.String("listen", nanoDefaultListen, "listen address for the hasJoined endpoint")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
@@ -61,24 +72,67 @@ func cmdNano(args []string, stdout, stderr io.Writer) int {
|
||||
fmt.Fprintln(stderr, "felis nano:", err)
|
||||
return 1
|
||||
}
|
||||
// [server] listen belongs to felis api. Someone moving nano off loopback naturally reaches
|
||||
// for it, and without this line would get connection refused with no hint why.
|
||||
if cfg.Server.Listen != "" {
|
||||
fmt.Fprintf(stderr, "felis nano: [server] listen = %q is ignored; nano binds -listen (%s), which the installer sets from FELIS_NANO_LISTEN\n", cfg.Server.Listen, *listen)
|
||||
}
|
||||
|
||||
handler := api.HasJoinedHandler(authSourcesFromConfig(cfg.AuthSources), nanoStubRepo{})
|
||||
fmt.Fprintf(stderr, "felis nano: hasJoined multiplexer on %s — Mojang + %d third-party source(s)\n", *listen, len(cfg.AuthSources))
|
||||
for i, s := range cfg.AuthSources {
|
||||
fmt.Fprintf(stderr, " [%d] %s -> %s\n", i+1, s.Tag, s.URL)
|
||||
}
|
||||
|
||||
// Log each request so a live login attempt is visible while testing against a real
|
||||
// Velocity — "is authlib even reaching me?" is the first question during verification.
|
||||
logged := http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
fmt.Fprintf(stderr, "felis nano: %s %s\n", r.Method, r.RequestURI)
|
||||
handler.ServeHTTP(w, r)
|
||||
})
|
||||
|
||||
srv := newAPIServer(*listen, logged)
|
||||
if err := srv.ListenAndServe(); err != nil {
|
||||
ln, err := net.Listen("tcp", *listen)
|
||||
if err != nil {
|
||||
fmt.Fprintln(stderr, "felis nano:", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
return serveNano(ctx, newAPIServer(*listen, nanoHandler(cfg.AuthSources, stderr)), ln, stderr)
|
||||
}
|
||||
|
||||
// nanoDrainTimeout outlasts the source scan of any realistic list (each source is given
|
||||
// five seconds) and stays well inside systemd's default 90-second stop timeout.
|
||||
const nanoDrainTimeout = 30 * time.Second
|
||||
|
||||
// serveNano serves until ctx ends, then drains. A restart, the documented way to pick up a
|
||||
// config edit, sends SIGTERM; without the drain a login already waiting on an upstream has
|
||||
// its connection reset, and Velocity tells that player the auth servers are down.
|
||||
func serveNano(ctx context.Context, srv *http.Server, ln net.Listener, stderr io.Writer) int {
|
||||
errc := make(chan error, 1)
|
||||
go func() { errc <- srv.Serve(ln) }()
|
||||
select {
|
||||
case err := <-errc:
|
||||
fmt.Fprintln(stderr, "felis nano:", err)
|
||||
return 1
|
||||
case <-ctx.Done():
|
||||
shutdownCtx, cancel := context.WithTimeout(context.Background(), nanoDrainTimeout)
|
||||
defer cancel()
|
||||
if err := srv.Shutdown(shutdownCtx); err != nil {
|
||||
fmt.Fprintln(stderr, "felis nano: shutdown:", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
}
|
||||
|
||||
// nanoHandler is what felis nano serves: the shared hasJoined handler, Mojang first, behind
|
||||
// a request log.
|
||||
func nanoHandler(sources []config.AuthSourceConfig, stderr io.Writer) http.Handler {
|
||||
handler := api.HasJoinedHandler(authSourcesFromConfig(sources), nanoStubRepo{})
|
||||
// Log each request so a live login attempt is visible while testing against a real
|
||||
// Velocity — "is Velocity even reaching me?" is the first question during verification.
|
||||
// The URI is the caller's text: quoted so a control or bidi character cannot rewrite the
|
||||
// line and invalid UTF-8 cannot turn the journal entry into a blob, and capped so one
|
||||
// request cannot write a megabyte of log.
|
||||
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
uri := r.RequestURI
|
||||
if len(uri) > nanoLogURIMax {
|
||||
uri = uri[:nanoLogURIMax] + "..."
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis nano: %s %q\n", r.Method, uri)
|
||||
handler.ServeHTTP(w, r)
|
||||
})
|
||||
}
|
||||
@@ -0,0 +1,139 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"io"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
"unicode/utf8"
|
||||
|
||||
"felis.lolicon.best/internal/api"
|
||||
)
|
||||
|
||||
// [server] listen in a nano config reads like the bind address but is not one; nano must
|
||||
// say so. The -listen value cannot be bound, so cmdNano returns right after loading.
|
||||
func TestNanoWarnsThatServerListenIsIgnored(t *testing.T) {
|
||||
cfg := filepath.Join(t.TempDir(), "felis.toml")
|
||||
if err := os.WriteFile(cfg, []byte("[server]\nlisten = \"0.0.0.0:9999\"\n"), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
var stderr bytes.Buffer
|
||||
if rc := cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr); rc != 1 {
|
||||
t.Fatalf("cmdNano = %d, want 1 from the unbindable -listen", rc)
|
||||
}
|
||||
if !strings.Contains(stderr.String(), `listen = "0.0.0.0:9999" is ignored`) {
|
||||
t.Fatalf("stderr %q should say the configured listen is ignored", stderr.String())
|
||||
}
|
||||
|
||||
// With no [server] table at all there is nothing to warn about.
|
||||
if err := os.WriteFile(cfg, nil, 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
stderr.Reset()
|
||||
_ = cmdNano([]string{"-config", cfg, "-listen", "127.0.0.1:-1"}, io.Discard, &stderr)
|
||||
if strings.Contains(stderr.String(), "is ignored") {
|
||||
t.Fatalf("stderr %q warns about a listen the operator never set", stderr.String())
|
||||
}
|
||||
}
|
||||
|
||||
// A stop signal that lands while a login is waiting on an upstream must let that login
|
||||
// finish: the request is answered, and serveNano returns only afterwards.
|
||||
func TestNanoDrainsInFlightLoginOnShutdown(t *testing.T) {
|
||||
entered, release := make(chan struct{}), make(chan struct{})
|
||||
srv := newAPIServer("", http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
close(entered)
|
||||
<-release
|
||||
w.WriteHeader(http.StatusNoContent)
|
||||
}))
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
ctx, stop := context.WithCancel(context.Background())
|
||||
done := make(chan int, 1)
|
||||
go func() { done <- serveNano(ctx, srv, ln, io.Discard) }()
|
||||
|
||||
got := make(chan int, 1)
|
||||
go func() {
|
||||
resp, err := http.Get("http://" + ln.Addr().String() + "/session/minecraft/hasJoined")
|
||||
if err != nil {
|
||||
got <- -1
|
||||
return
|
||||
}
|
||||
resp.Body.Close()
|
||||
got <- resp.StatusCode
|
||||
}()
|
||||
<-entered
|
||||
stop()
|
||||
select {
|
||||
case <-done:
|
||||
t.Fatal("serveNano returned while a login was still in flight")
|
||||
case <-time.After(200 * time.Millisecond):
|
||||
}
|
||||
close(release)
|
||||
if code := <-got; code != http.StatusNoContent {
|
||||
t.Fatalf("in-flight login got %d, want its answer (204)", code)
|
||||
}
|
||||
if rc := <-done; rc != 0 {
|
||||
t.Fatalf("serveNano = %d after a clean drain, want 0", rc)
|
||||
}
|
||||
}
|
||||
|
||||
// The nano delivery path: the shared handler behind nano's stub store must admit a login its
|
||||
// source validated. nanoStubRepo implements only the bar-list lookup, so a new store call in
|
||||
// handleHasJoined would reach its nil embedded Repo and panic here, while the full-api tests,
|
||||
// which use a complete fake store, stay green.
|
||||
func TestNanoAdmitsAValidatedLogin(t *testing.T) {
|
||||
const id = "069a79f444e94726a5befca90e38aaf5"
|
||||
ygg := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
_, _ = io.WriteString(w, `{"id":"`+id+`","name":"Notch"}`)
|
||||
}))
|
||||
defer ygg.Close()
|
||||
h := api.HasJoinedHandler([]api.AuthSource{{Tag: "mojang", URL: ygg.URL, Identity: true}}, nanoStubRepo{})
|
||||
w := httptest.NewRecorder()
|
||||
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, "/session/minecraft/hasJoined?username=Notch&serverId=abc", nil))
|
||||
if w.Code != http.StatusOK || !strings.Contains(w.Body.String(), id) {
|
||||
t.Fatalf("code = %d body = %q, want the validated profile", w.Code, w.Body.String())
|
||||
}
|
||||
}
|
||||
|
||||
// An unauthenticated relay on a public address spends this host's Mojang rate limit for
|
||||
// anyone who finds it, so the default bind has to stay loopback.
|
||||
func TestNanoListensOnLoopbackByDefault(t *testing.T) {
|
||||
host, _, err := net.SplitHostPort(nanoDefaultListen)
|
||||
if ip := net.ParseIP(host); err != nil || ip == nil || !ip.IsLoopback() {
|
||||
t.Fatalf("default -listen %q is not a loopback address", nanoDefaultListen)
|
||||
}
|
||||
}
|
||||
|
||||
// The request log prints text the caller chose. A bidi override must not reorder the line,
|
||||
// an invalid byte must not make journald store the entry as a blob, and a huge query must
|
||||
// not become a huge log line. serverId is left out so the handler answers without asking
|
||||
// any source.
|
||||
func TestNanoRequestLogIsQuotedAndCapped(t *testing.T) {
|
||||
const rlo = rune(0x202e) // RIGHT-TO-LEFT OVERRIDE
|
||||
var log bytes.Buffer
|
||||
h := nanoHandler(nil, &log)
|
||||
target := "/session/minecraft/hasJoined?username=" + string(rlo) + "evil" + string([]byte{0x9b}) + "31m" + strings.Repeat("a", 4096)
|
||||
w := httptest.NewRecorder()
|
||||
h.ServeHTTP(w, httptest.NewRequest(http.MethodGet, target, nil))
|
||||
|
||||
line := log.String()
|
||||
if strings.ContainsRune(line, rlo) || !utf8.ValidString(line) {
|
||||
t.Fatalf("raw caller bytes reached the log: %q", line)
|
||||
}
|
||||
if escaped := strings.Trim(strconv.QuoteRune(rlo), "'"); !strings.Contains(line, escaped) {
|
||||
t.Fatalf("log line %q should show the override escaped as %s", line, escaped)
|
||||
}
|
||||
if len(line) > 2*nanoLogURIMax {
|
||||
t.Fatalf("log line is %d bytes for a %d-byte URI; want it capped", len(line), len(target))
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,556 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bufio"
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/dbbackup"
|
||||
"felis.lolicon.best/internal/offsite"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"felis.lolicon.best/internal/store"
|
||||
appsv1 "k8s.io/api/apps/v1"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"k8s.io/apimachinery/pkg/types"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
)
|
||||
|
||||
const offsiteUsage = `usage:
|
||||
felis offsite sync [-config path] [-archive-dir dir] [-db-dir dir] [-status-file path]
|
||||
felis offsite status [-config path] [-status-file path]
|
||||
felis offsite list [-config path]
|
||||
felis offsite fetch-db [-config path | -endpoint url -bucket name [-region r] [-prefix p]]
|
||||
[-dir dir] latest|<bundle>
|
||||
felis offsite fetch-worlds [-config path] [-archive-dir dir]
|
||||
felis offsite keygen
|
||||
|
||||
Every verb but keygen reads the bucket credentials and the encryption key from
|
||||
the variables [offsite] names (default FELIS_OFFSITE_ACCESS_KEY,
|
||||
FELIS_OFFSITE_SECRET_KEY, FELIS_OFFSITE_KEY), taking any that are unset from
|
||||
-env-file (default /etc/felis/offsite.env).
|
||||
`
|
||||
|
||||
// defaultOffsiteEnvFile is where bootstrap keeps the [offsite] secrets; the
|
||||
// felis-offsite.service unit loads it as its EnvironmentFile.
|
||||
const defaultOffsiteEnvFile = "/etc/felis/offsite.env"
|
||||
|
||||
// cmdOffsite implements `felis offsite`: the off-site copy of the world
|
||||
// archives and the database bundles (internal/offsite). felis-offsite.timer
|
||||
// runs `sync` hourly on the host; the fetch verbs are the way back after the
|
||||
// node is lost (docs/troubleshooting.md §16).
|
||||
func cmdOffsite(args []string, stdout, stderr io.Writer) int {
|
||||
if len(args) == 0 {
|
||||
fmt.Fprint(stderr, offsiteUsage)
|
||||
return 2
|
||||
}
|
||||
verb, rest := args[0], args[1:]
|
||||
fs := flag.NewFlagSet("offsite "+verb, flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
fs.Usage = func() { fmt.Fprint(stderr, offsiteUsage) }
|
||||
switch verb {
|
||||
case "sync":
|
||||
return offsiteSync(fs, rest, stdout, stderr)
|
||||
case "status":
|
||||
return offsiteStatus(fs, rest, stdout, stderr)
|
||||
case "list":
|
||||
return offsiteList(fs, rest, stdout, stderr)
|
||||
case "fetch-db":
|
||||
return offsiteFetchDB(fs, rest, stdout, stderr)
|
||||
case "fetch-worlds":
|
||||
return offsiteFetchWorlds(fs, rest, stdout, stderr)
|
||||
case "keygen":
|
||||
k, err := offsite.NewKey()
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite keygen: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintln(stdout, k)
|
||||
return 0
|
||||
case "-h", "--help", "help":
|
||||
fmt.Fprint(stdout, offsiteUsage)
|
||||
return 0
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis offsite: unknown verb %q\n%s", verb, offsiteUsage)
|
||||
return 2
|
||||
}
|
||||
|
||||
// offsiteEnv is the resolved [offsite] binding: the bucket and the key.
|
||||
type offsiteEnv struct {
|
||||
cfg config.OffsiteConfig
|
||||
bucket *offsite.S3
|
||||
key []byte
|
||||
}
|
||||
|
||||
// loadEnvFile sets each KEY=VALUE of path that is not already in the
|
||||
// environment, so a root shell runs a command the same way its unit does.
|
||||
// A missing file is not an error.
|
||||
func loadEnvFile(path string) error {
|
||||
if path == "" {
|
||||
return nil
|
||||
}
|
||||
f, err := os.Open(path)
|
||||
if errors.Is(err, os.ErrNotExist) {
|
||||
return nil
|
||||
}
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
defer f.Close()
|
||||
sc := bufio.NewScanner(f)
|
||||
for sc.Scan() {
|
||||
line := strings.TrimSpace(sc.Text())
|
||||
if line == "" || strings.HasPrefix(line, "#") {
|
||||
continue
|
||||
}
|
||||
k, v, ok := strings.Cut(line, "=")
|
||||
if !ok {
|
||||
continue
|
||||
}
|
||||
k = strings.TrimSpace(strings.TrimPrefix(k, "export "))
|
||||
v = strings.TrimSpace(v)
|
||||
if len(v) >= 2 && (v[0] == '"' || v[0] == '\'') && v[len(v)-1] == v[0] {
|
||||
v = v[1 : len(v)-1]
|
||||
}
|
||||
if os.Getenv(k) == "" {
|
||||
os.Setenv(k, v)
|
||||
}
|
||||
}
|
||||
return sc.Err()
|
||||
}
|
||||
|
||||
// resolveOffsite builds the bucket client and parses the key for c.
|
||||
func resolveOffsite(c config.OffsiteConfig) (*offsiteEnv, error) {
|
||||
if !c.Enabled() {
|
||||
return nil, errors.New("no [offsite] bucket is configured (docs/troubleshooting.md §16, \"Keep a copy somewhere else\")")
|
||||
}
|
||||
need := func(ref, what string) (string, error) {
|
||||
v := os.Getenv(ref)
|
||||
if v == "" {
|
||||
return "", fmt.Errorf("%s: environment variable %s is empty (set it, or put it in %s)", what, ref, defaultOffsiteEnvFile)
|
||||
}
|
||||
return v, nil
|
||||
}
|
||||
ak, err := need(c.AccessKeyRef, "access key")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
sk, err := need(c.SecretKeyRef, "secret key")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
rawKey, err := need(c.KeyRef, "encryption key")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
key, err := offsite.ParseKey(rawKey)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
b, err := offsite.NewS3(offsite.S3Config{
|
||||
Endpoint: c.Endpoint, Region: c.Region, Bucket: c.Bucket, Prefix: c.Prefix,
|
||||
AccessKey: ak, SecretKey: sk,
|
||||
})
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
return &offsiteEnv{cfg: c, bucket: b, key: key}, nil
|
||||
}
|
||||
|
||||
// loadOffsite loads felis.toml and the env file and resolves [offsite].
|
||||
func loadOffsite(cfgPath, envFile string) (*config.Config, *offsiteEnv, error) {
|
||||
if err := loadEnvFile(envFile); err != nil {
|
||||
return nil, nil, fmt.Errorf("read %s: %w", envFile, err)
|
||||
}
|
||||
cfg, err := config.Load(cfgPath)
|
||||
if err != nil {
|
||||
return nil, nil, err
|
||||
}
|
||||
env, err := resolveOffsite(cfg.Offsite)
|
||||
if err != nil {
|
||||
return cfg, nil, err
|
||||
}
|
||||
return cfg, env, nil
|
||||
}
|
||||
|
||||
func offsiteSync(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy, which reaches PostgreSQL on 127.0.0.1)")
|
||||
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
|
||||
archiveDir := fs.String("archive-dir", "", "host directory of the world archive volume (default: resolved from the backup PVC through the cluster)")
|
||||
backupPVC := fs.String("backup-pvc", "felis-backups", `the world archive PVC, in the [k8s] namespace ("" when backups are off)`)
|
||||
dbDir := fs.String("db-dir", dbbackup.DefaultDir, `database bundle directory ("" copies no bundles)`)
|
||||
statusFile := fs.String("status-file", offsite.DefaultStatusFile, "where the result of this run is recorded for the watchdog and `status`")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
cfg, env, err := loadOffsite(*cfgPath, *envFile)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite sync: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
st := offsite.Status{
|
||||
LastAttempt: time.Now().UTC(), Endpoint: env.cfg.Endpoint, Bucket: env.cfg.Bucket,
|
||||
Prefix: env.cfg.Prefix, KeyID: offsite.KeyID(env.key),
|
||||
}
|
||||
if prev, _ := offsite.ReadStatus(*statusFile); prev != nil {
|
||||
st.LastSuccess = prev.LastSuccess
|
||||
}
|
||||
res, err := runOffsiteSync(cfg, env, *archiveDir, *backupPVC, *dbDir, stderr)
|
||||
st.Result = res
|
||||
if err != nil {
|
||||
st.LastError = err.Error()
|
||||
} else {
|
||||
st.LastSuccess = st.LastAttempt
|
||||
}
|
||||
if werr := offsite.WriteStatus(*statusFile, st); werr != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite sync: record status: %v\n", werr)
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis offsite sync: worlds copied=%d pending=%d missing=%d expired=%d; bundles copied=%d pruned=%d; bucket holds %d worlds (%s) and %d bundles\n",
|
||||
res.WorldsUploaded, res.WorldsPending, len(res.WorldsMissing), res.WorldsExpired,
|
||||
res.DBUploaded, res.DBPruned, res.RemoteWorlds, offsite.HumanBytes(res.RemoteBytes), res.RemoteDB)
|
||||
for _, m := range res.WorldsMissing {
|
||||
fmt.Fprintf(stderr, "felis offsite sync: recorded archive not on the volume, nothing to copy: %s\n", m)
|
||||
}
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite sync: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func runOffsiteSync(cfg *config.Config, env *offsiteEnv, archiveDir, backupPVC, dbDir string, log io.Writer) (offsite.Result, error) {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 50*time.Minute)
|
||||
defer cancel()
|
||||
checkCtx, checkCancel := context.WithTimeout(ctx, 30*time.Second)
|
||||
err := env.bucket.Check(checkCtx)
|
||||
checkCancel()
|
||||
if err != nil {
|
||||
return offsite.Result{}, err
|
||||
}
|
||||
if archiveDir == "" && backupPVC != "" {
|
||||
dir, err := resolveArchiveDir(ctx, cfg.K8s.Namespace, backupPVC, false, log)
|
||||
if err != nil {
|
||||
return offsite.Result{}, err
|
||||
}
|
||||
archiveDir = dir
|
||||
}
|
||||
drv, err := store.Open(ctx, cfg.Database.URL)
|
||||
if err != nil {
|
||||
return offsite.Result{}, fmt.Errorf("open database: %w", err)
|
||||
}
|
||||
defer drv.Close()
|
||||
s := &offsite.Syncer{
|
||||
Bucket: env.bucket, Catalog: offsite.PGCatalog{DB: drv.DB()}, Key: env.key,
|
||||
ArchiveDir: archiveDir, DBDir: dbDir, DBKeep: env.cfg.DBKeep, Log: log,
|
||||
}
|
||||
return s.Run(ctx)
|
||||
}
|
||||
|
||||
// resolveArchiveDir finds the host directory behind the world archive PVC: a
|
||||
// local-path volume is a directory on this node. A PVC still waiting for its
|
||||
// first consumer holds nothing yet: without bind that is "" (no archives),
|
||||
// with bind it is bound first, for fetch-worlds to write into.
|
||||
func resolveArchiveDir(ctx context.Context, ns, pvcName string, bind bool, log io.Writer) (string, error) {
|
||||
if ns == "" {
|
||||
ns = platform.DefaultMinecraftNamespace
|
||||
}
|
||||
cl, err := buildSystemServerClient()
|
||||
if err != nil {
|
||||
return "", fmt.Errorf("reach the cluster to find the archive volume (or pass -archive-dir): %w", err)
|
||||
}
|
||||
var pvc corev1.PersistentVolumeClaim
|
||||
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err != nil {
|
||||
return "", fmt.Errorf("archive volume %s/%s: %w", ns, pvcName, err)
|
||||
}
|
||||
if pvc.Spec.VolumeName == "" {
|
||||
if !bind {
|
||||
fmt.Fprintf(log, "felis offsite: archive volume %s/%s is not bound yet; no world has been archived\n", ns, pvcName)
|
||||
return "", nil
|
||||
}
|
||||
if err := bindVolume(ctx, cl, ns, pvcName, log); err != nil {
|
||||
return "", err
|
||||
}
|
||||
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err != nil {
|
||||
return "", err
|
||||
}
|
||||
}
|
||||
var pv corev1.PersistentVolume
|
||||
if err := cl.Get(ctx, types.NamespacedName{Name: pvc.Spec.VolumeName}, &pv); err != nil {
|
||||
return "", fmt.Errorf("archive volume %s: %w", pvc.Spec.VolumeName, err)
|
||||
}
|
||||
var dir string
|
||||
switch {
|
||||
case pv.Spec.Local != nil:
|
||||
dir = pv.Spec.Local.Path
|
||||
case pv.Spec.HostPath != nil:
|
||||
dir = pv.Spec.HostPath.Path
|
||||
default:
|
||||
return "", fmt.Errorf("archive volume %s is not a directory on a node (local or hostPath); pass -archive-dir with where it is mounted on this host", pv.Name)
|
||||
}
|
||||
if fi, err := os.Stat(dir); err != nil || !fi.IsDir() {
|
||||
return "", fmt.Errorf("archive volume %s is %s on its node, which is not a directory here; run this on the node that holds it, or pass -archive-dir", pv.Name, dir)
|
||||
}
|
||||
return dir, nil
|
||||
}
|
||||
|
||||
// bindVolume runs a pod that mounts the PVC and exits, which is what makes a
|
||||
// WaitForFirstConsumer volume (k3s local-path) get provisioned. The pod uses
|
||||
// the control plane's own image, which every install already has.
|
||||
func bindVolume(ctx context.Context, cl client.Client, ns, pvcName string, log io.Writer) error {
|
||||
var api appsv1.Deployment
|
||||
if err := cl.Get(ctx, types.NamespacedName{Namespace: platform.DefaultControlNamespace, Name: "felis-api"}, &api); err != nil {
|
||||
return fmt.Errorf("find the felis image to bind the archive volume with: %w", err)
|
||||
}
|
||||
if len(api.Spec.Template.Spec.Containers) == 0 {
|
||||
return errors.New("felis-api has no container to take the image from")
|
||||
}
|
||||
image := api.Spec.Template.Spec.Containers[0].Image
|
||||
pod := platform.VolumeBinderPod(ns, pvcName, image)
|
||||
if err := cl.Create(ctx, pod); err != nil {
|
||||
return fmt.Errorf("start a pod to bind the archive volume: %w", err)
|
||||
}
|
||||
fmt.Fprintf(log, "felis offsite: binding the archive volume %s/%s (pod %s)\n", ns, pvcName, pod.Name)
|
||||
defer func() {
|
||||
_ = cl.Delete(context.Background(), pod, client.PropagationPolicy(metav1.DeletePropagationBackground))
|
||||
}()
|
||||
deadline := time.Now().Add(3 * time.Minute)
|
||||
for time.Now().Before(deadline) {
|
||||
var pvc corev1.PersistentVolumeClaim
|
||||
if err := cl.Get(ctx, types.NamespacedName{Namespace: ns, Name: pvcName}, &pvc); err == nil && pvc.Spec.VolumeName != "" && pvc.Status.Phase == corev1.ClaimBound {
|
||||
return nil
|
||||
}
|
||||
select {
|
||||
case <-ctx.Done():
|
||||
return ctx.Err()
|
||||
case <-time.After(2 * time.Second):
|
||||
}
|
||||
}
|
||||
return fmt.Errorf("the archive volume %s/%s did not bind within 3 minutes; see kubectl -n %s describe pod %s", ns, pvcName, ns, pod.Name)
|
||||
}
|
||||
|
||||
func offsiteStatus(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
statusFile := fs.String("status-file", offsite.DefaultStatusFile, "the record `sync` writes")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite status: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if !cfg.Offsite.Enabled() {
|
||||
fmt.Fprintln(stdout, "off-site copy: not configured. World archives and database bundles exist on this machine only.")
|
||||
fmt.Fprintln(stdout, "See docs/troubleshooting.md §16, \"Keep a copy somewhere else\".")
|
||||
return 1
|
||||
}
|
||||
o := cfg.Offsite
|
||||
fmt.Fprintf(stdout, "bucket: %s at %s", o.Bucket, o.Endpoint)
|
||||
if o.Prefix != "" {
|
||||
fmt.Fprintf(stdout, ", prefix %s", o.Prefix)
|
||||
}
|
||||
fmt.Fprintln(stdout)
|
||||
st, err := offsite.ReadStatus(*statusFile)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite status: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if st == nil {
|
||||
fmt.Fprintln(stdout, "last sync: never (sudo systemctl start felis-offsite.service)")
|
||||
return 1
|
||||
}
|
||||
now := time.Now()
|
||||
fmt.Fprintf(stdout, "key id: %s\n", st.KeyID)
|
||||
fmt.Fprintf(stdout, "last attempt: %s (%s ago)\n", st.LastAttempt.Local().Format(time.DateTime), dbbackup.Age(now.Sub(st.LastAttempt)))
|
||||
if st.LastSuccess.IsZero() {
|
||||
fmt.Fprintln(stdout, "last success: never")
|
||||
} else {
|
||||
fmt.Fprintf(stdout, "last success: %s (%s ago)\n", st.LastSuccess.Local().Format(time.DateTime), dbbackup.Age(now.Sub(st.LastSuccess)))
|
||||
}
|
||||
if st.LastError != "" {
|
||||
fmt.Fprintf(stdout, "last error: %s\n", st.LastError)
|
||||
}
|
||||
r := st.Result
|
||||
fmt.Fprintf(stdout, "bucket holds: %d world archives (%s), %d database bundles, newest %s\n",
|
||||
r.RemoteWorlds, offsite.HumanBytes(r.RemoteBytes), r.RemoteDB, orNone(r.NewestDB))
|
||||
fmt.Fprintf(stdout, "waiting: %d world archives not yet copied\n", r.WorldsPending)
|
||||
for _, m := range r.WorldsMissing {
|
||||
fmt.Fprintf(stdout, "missing: %s is recorded but not on the volume\n", m)
|
||||
}
|
||||
if st.LastSuccess.IsZero() || now.Sub(st.LastSuccess) > offsite.StaleAfter {
|
||||
fmt.Fprintf(stdout, "\nThe last successful sync is older than %s: journalctl -u felis-offsite -n 50\n", dbbackup.Age(offsite.StaleAfter))
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
func orNone(s string) string {
|
||||
if s == "" {
|
||||
return "none"
|
||||
}
|
||||
return s
|
||||
}
|
||||
|
||||
func offsiteList(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
_, env, err := loadOffsite(*cfgPath, *envFile)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
return printOffsiteList(env, stdout, stderr)
|
||||
}
|
||||
|
||||
func printOffsiteList(env *offsiteEnv, stdout, stderr io.Writer) int {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 2*time.Minute)
|
||||
defer cancel()
|
||||
bundles, err := offsite.ListDB(ctx, env.bucket)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
worlds, err := env.bucket.List(ctx, "worlds/")
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite list: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "database bundles (%d, newest first):\n", len(bundles))
|
||||
for _, b := range bundles {
|
||||
fmt.Fprintf(stdout, " %s %s\n", b.Key, offsite.HumanBytes(b.Size))
|
||||
}
|
||||
var total int64
|
||||
for _, w := range worlds {
|
||||
total += w.Size
|
||||
}
|
||||
fmt.Fprintf(stdout, "world archives: %d (%s)\n", len(worlds), offsite.HumanBytes(total))
|
||||
return 0
|
||||
}
|
||||
|
||||
func offsiteFetchDB(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml; on a host with no install yet, give -endpoint and -bucket instead")
|
||||
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
|
||||
endpoint := fs.String("endpoint", "", "bucket endpoint, when there is no felis.toml")
|
||||
bucket := fs.String("bucket", "", "bucket name, when there is no felis.toml")
|
||||
region := fs.String("region", "", "bucket region, when there is no felis.toml")
|
||||
prefix := fs.String("prefix", "", "key prefix, when there is no felis.toml")
|
||||
dir := fs.String("dir", dbbackup.DefaultDir, "directory to write the bundle to")
|
||||
arg, ok := parseWithArg(fs, args)
|
||||
if !ok {
|
||||
return 2
|
||||
}
|
||||
if arg == "" {
|
||||
fmt.Fprint(stderr, offsiteUsage)
|
||||
return 2
|
||||
}
|
||||
if err := loadEnvFile(*envFile); err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: read %s: %v\n", *envFile, err)
|
||||
return 1
|
||||
}
|
||||
var oc config.OffsiteConfig
|
||||
if *bucket != "" {
|
||||
oc = config.OffsiteConfig{
|
||||
Endpoint: *endpoint, Bucket: *bucket, Region: *region, Prefix: *prefix,
|
||||
AccessKeyRef: config.DefaultOffsiteAccessKeyEnv, SecretKeyRef: config.DefaultOffsiteSecretKeyEnv,
|
||||
KeyRef: config.DefaultOffsiteKeyEnv,
|
||||
}
|
||||
} else {
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %v (on a host with no install yet, pass -endpoint and -bucket)\n", err)
|
||||
return 1
|
||||
}
|
||||
oc = cfg.Offsite
|
||||
}
|
||||
env, err := resolveOffsite(oc)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Minute)
|
||||
defer cancel()
|
||||
name := arg
|
||||
if name == "latest" {
|
||||
bundles, err := offsite.ListDB(ctx, env.bucket)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if len(bundles) == 0 {
|
||||
fmt.Fprintln(stderr, "felis offsite fetch-db: the bucket holds no database bundle")
|
||||
return 1
|
||||
}
|
||||
name = bundles[0].Key
|
||||
}
|
||||
if _, _, ok := dbbackup.ParseBundleName(name); !ok {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %q is not a bundle name (felis-db-<stamp>-<label>.tar); see `felis offsite list`\n", name)
|
||||
return 2
|
||||
}
|
||||
if err := os.MkdirAll(*dir, 0o700); err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
dst := filepath.Join(*dir, name)
|
||||
if err := offsite.FetchObject(ctx, env.bucket, env.key, offsite.DBKey(name), dst, 0o600); err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if _, err := dbbackup.Verify(dst); err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-db: fetched %s but it does not verify: %v\n", dst, err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis offsite fetch-db: wrote %s (verified)\n", dst)
|
||||
return 0
|
||||
}
|
||||
|
||||
func offsiteFetchWorlds(fs *flag.FlagSet, args []string, stdout, stderr io.Writer) int {
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy)")
|
||||
envFile := fs.String("env-file", defaultOffsiteEnvFile, "file with the [offsite] secrets, for variables not already set")
|
||||
archiveDir := fs.String("archive-dir", "", "host directory of the world archive volume (default: resolved from the backup PVC, binding it if needed)")
|
||||
backupPVC := fs.String("backup-pvc", "felis-backups", "the world archive PVC, in the [k8s] namespace")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
cfg, env, err := loadOffsite(*cfgPath, *envFile)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 6*time.Hour)
|
||||
defer cancel()
|
||||
dir := *archiveDir
|
||||
if dir == "" {
|
||||
if dir, err = resolveArchiveDir(ctx, cfg.K8s.Namespace, *backupPVC, true, stderr); err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
}
|
||||
drv, err := store.Open(ctx, cfg.Database.URL)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-worlds: open database: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer drv.Close()
|
||||
res, err := offsite.FetchWorlds(ctx, env.bucket, offsite.PGCatalog{DB: drv.DB()}, env.key, dir, stderr)
|
||||
fmt.Fprintf(stdout, "felis offsite fetch-worlds: %d recorded archives, %d fetched into %s, %d with no copy in the bucket\n",
|
||||
res.Present, len(res.Fetched), dir, len(res.Missing))
|
||||
for _, m := range res.Missing {
|
||||
fmt.Fprintf(stdout, " no off-site copy: %s\n", m)
|
||||
}
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis offsite fetch-worlds: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
@@ -0,0 +1,96 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/offsite"
|
||||
)
|
||||
|
||||
func TestLoadOffsiteEnvFile(t *testing.T) {
|
||||
path := filepath.Join(t.TempDir(), "offsite.env")
|
||||
body := `# written by bootstrap
|
||||
FELIS_OFFSITE_ACCESS_KEY=AKIA123
|
||||
export FELIS_OFFSITE_SECRET_KEY="se=cret"
|
||||
FELIS_OFFSITE_KEY='k'
|
||||
|
||||
not a line
|
||||
`
|
||||
if err := os.WriteFile(path, []byte(body), 0o600); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
t.Setenv("FELIS_OFFSITE_ACCESS_KEY", "from-the-shell")
|
||||
t.Setenv("FELIS_OFFSITE_SECRET_KEY", "")
|
||||
t.Setenv("FELIS_OFFSITE_KEY", "")
|
||||
if err := loadEnvFile(path); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
for k, want := range map[string]string{
|
||||
"FELIS_OFFSITE_ACCESS_KEY": "from-the-shell", // the environment wins
|
||||
"FELIS_OFFSITE_SECRET_KEY": "se=cret",
|
||||
"FELIS_OFFSITE_KEY": "k",
|
||||
} {
|
||||
if got := os.Getenv(k); got != want {
|
||||
t.Errorf("%s = %q, want %q", k, got, want)
|
||||
}
|
||||
}
|
||||
if err := loadEnvFile(filepath.Join(t.TempDir(), "absent")); err != nil {
|
||||
t.Errorf("a missing env file is not an error: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveOffsiteNamesTheMissingVariable(t *testing.T) {
|
||||
c := config.OffsiteConfig{
|
||||
Endpoint: "https://s3.example", Bucket: "b",
|
||||
AccessKeyRef: "T_AK", SecretKeyRef: "T_SK", KeyRef: "T_KEY",
|
||||
}
|
||||
t.Setenv("T_AK", "ak")
|
||||
t.Setenv("T_SK", "sk")
|
||||
t.Setenv("T_KEY", "")
|
||||
if _, err := resolveOffsite(c); err == nil || !strings.Contains(err.Error(), "T_KEY") {
|
||||
t.Fatalf("err = %v, want it to name T_KEY", err)
|
||||
}
|
||||
t.Setenv("T_KEY", "not base64 at all")
|
||||
if _, err := resolveOffsite(c); err == nil {
|
||||
t.Fatal("a malformed key was accepted")
|
||||
}
|
||||
key, _ := offsite.NewKey()
|
||||
t.Setenv("T_KEY", key)
|
||||
env, err := resolveOffsite(c)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if len(env.key) != offsite.KeySize {
|
||||
t.Fatalf("key is %d bytes", len(env.key))
|
||||
}
|
||||
if _, err := resolveOffsite(config.OffsiteConfig{}); err == nil {
|
||||
t.Fatal("an unconfigured [offsite] resolved")
|
||||
}
|
||||
}
|
||||
|
||||
func TestOffsiteKeygen(t *testing.T) {
|
||||
var out, errb bytes.Buffer
|
||||
if code := cmdOffsite([]string{"keygen"}, &out, &errb); code != 0 {
|
||||
t.Fatalf("exit %d: %s", code, errb.String())
|
||||
}
|
||||
if _, err := offsite.ParseKey(strings.TrimSpace(out.String())); err != nil {
|
||||
t.Fatalf("keygen printed %q: %v", out.String(), err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOffsiteFetchDBRejectsOddNames(t *testing.T) {
|
||||
key, _ := offsite.NewKey()
|
||||
t.Setenv("FELIS_OFFSITE_ACCESS_KEY", "ak")
|
||||
t.Setenv("FELIS_OFFSITE_SECRET_KEY", "sk")
|
||||
t.Setenv("FELIS_OFFSITE_KEY", key)
|
||||
var out, errb bytes.Buffer
|
||||
code := cmdOffsite([]string{"fetch-db", "-env-file", "", "-endpoint", "http://127.0.0.1:1", "-bucket", "b",
|
||||
"-dir", t.TempDir(), "../../etc/shadow"}, &out, &errb)
|
||||
if code != 2 || !strings.Contains(errb.String(), "not a bundle name") {
|
||||
t.Fatalf("exit %d: %s", code, errb.String())
|
||||
}
|
||||
}
|
||||
+58
-2
@@ -1,19 +1,26 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net/http"
|
||||
"os"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
felismetrics "felis.lolicon.best/internal/metrics"
|
||||
"felis.lolicon.best/internal/operator"
|
||||
"github.com/go-logr/logr"
|
||||
"k8s.io/apimachinery/pkg/runtime"
|
||||
utilruntime "k8s.io/apimachinery/pkg/util/runtime"
|
||||
clientgoscheme "k8s.io/client-go/kubernetes/scheme"
|
||||
ctrl "sigs.k8s.io/controller-runtime"
|
||||
"sigs.k8s.io/controller-runtime/pkg/cache"
|
||||
"sigs.k8s.io/controller-runtime/pkg/healthz"
|
||||
ctrlmetrics "sigs.k8s.io/controller-runtime/pkg/metrics"
|
||||
metricsserver "sigs.k8s.io/controller-runtime/pkg/metrics/server"
|
||||
)
|
||||
@@ -25,6 +32,11 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("operator", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
metricsAddr := fs.String("metrics-bind-address", ":8080", "address the metric endpoint binds to")
|
||||
// healthAddr serves the manager's health endpoints (/healthz, /readyz) that the
|
||||
// Deployment's probes dial. Without it the operator pod would carry no probe at
|
||||
// all, and a wedged manager would keep its endpoint forever. It must differ from
|
||||
// metricsAddr: the metrics server owns :8080.
|
||||
healthAddr := fs.String("health-probe-bind-address", ":8081", "address the health probe endpoint binds to")
|
||||
// namespace MUST equal the [k8s] namespace felis-api is configured with, and
|
||||
// the deployment manifests (felis manifests) render both from one value. It
|
||||
// scopes the manager's cache (informers) to a single namespace so the operator
|
||||
@@ -42,9 +54,16 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
utilruntime.Must(clientgoscheme.AddToScheme(scheme))
|
||||
utilruntime.Must(v1alpha1.AddToScheme(scheme))
|
||||
|
||||
// controller-runtime logs through its own logr sink; without one, its first
|
||||
// reconcile prints "log.SetLogger(...) was never called" ATTACHED TO A FULL
|
||||
// GOROUTINE STACK — pure noise, not signal. Route it to slog's default handler
|
||||
// so its messages appear as ordinary stderr lines.
|
||||
ctrl.SetLogger(logr.FromSlogHandler(slog.Default().Handler()))
|
||||
|
||||
mgr, err := ctrl.NewManager(ctrl.GetConfigOrDie(), ctrl.Options{
|
||||
Scheme: scheme,
|
||||
Metrics: metricsserver.Options{BindAddress: *metricsAddr},
|
||||
Scheme: scheme,
|
||||
Metrics: metricsserver.Options{BindAddress: *metricsAddr},
|
||||
HealthProbeBindAddress: *healthAddr,
|
||||
// Scope every informer to the single watched namespace. Without this the
|
||||
// cached client (mgr.GetClient) would LIST/WATCH cluster-wide, which a
|
||||
// namespaced Role cannot grant — the operator would fail closed at runtime
|
||||
@@ -60,6 +79,23 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis operator: watching namespace %q\n", *namespace)
|
||||
|
||||
// /healthz fails while a reconcile pass has been stuck past its limit, so the
|
||||
// liveness probe restarts an operator whose workers are wedged (a Pod whose
|
||||
// process answers but no server starts or stops). /readyz waits for the
|
||||
// informer caches: until they sync the operator acts on nothing, and one that
|
||||
// never syncs (lost RBAC, an unreachable API) never reports Available.
|
||||
// A dependency hiccup fails neither: the caches ride through API blips, and
|
||||
// each pass is bounded well inside the stuck limit.
|
||||
watch := &operator.ReconcileWatch{}
|
||||
if err := mgr.AddHealthzCheck("reconcile", watch.Check); err != nil {
|
||||
fmt.Fprintf(stderr, "felis operator: register healthz check: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if err := mgr.AddReadyzCheck("informers", cacheSynced(mgr.GetCache())); err != nil {
|
||||
fmt.Fprintf(stderr, "felis operator: register readyz check: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
// Publish the named felis_* metrics (spec §23) on the endpoint the manager
|
||||
// already serves (metricsAddr). controller-runtime's metrics server exposes
|
||||
// its global Registry, so registering into it is all that is needed for
|
||||
@@ -69,6 +105,7 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
fmt.Fprintf(stderr, "felis operator: register metrics: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
felismetrics.SetBuildInfo("operator", resolvedVersion())
|
||||
|
||||
r := &operator.Reconciler{
|
||||
Client: mgr.GetClient(),
|
||||
@@ -78,6 +115,10 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
// injects into user servers. The Deployment passes it as FELIS_IMAGE (see
|
||||
// platform.OperatorDeployment); absent, that injection is simply skipped.
|
||||
FelisImage: os.Getenv("FELIS_IMAGE"),
|
||||
// Uncached: the maintenance-lock check lists Jobs only when a server is
|
||||
// about to start, which does not justify a namespace-wide Job informer.
|
||||
Jobs: mgr.GetAPIReader(),
|
||||
Watch: watch,
|
||||
}
|
||||
if err := r.SetupWithManager(mgr); err != nil {
|
||||
fmt.Fprintf(stderr, "felis operator: setup controller: %v\n", err)
|
||||
@@ -99,3 +140,18 @@ func cmdOperator(args []string, _, stderr io.Writer) int {
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// cacheSynced is a readyz check that passes once every informer the manager
|
||||
// started has synced. It waits at most a second, well inside the probe timeout.
|
||||
func cacheSynced(c interface {
|
||||
WaitForCacheSync(ctx context.Context) bool
|
||||
}) healthz.Checker {
|
||||
return func(req *http.Request) error {
|
||||
ctx, cancel := context.WithTimeout(req.Context(), time.Second)
|
||||
defer cancel()
|
||||
if !c.WaitForCacheSync(ctx) {
|
||||
return errors.New("informer caches not synced")
|
||||
}
|
||||
return nil
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,28 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
)
|
||||
|
||||
type fakeCache bool
|
||||
|
||||
func (f fakeCache) WaitForCacheSync(ctx context.Context) bool {
|
||||
if !f {
|
||||
<-ctx.Done()
|
||||
}
|
||||
return bool(f)
|
||||
}
|
||||
|
||||
// TestCacheSynced: the operator reports ready only once its informers synced,
|
||||
// and a check against caches that never sync returns within its own deadline.
|
||||
func TestCacheSynced(t *testing.T) {
|
||||
req := httptest.NewRequest("GET", "/readyz", nil)
|
||||
if err := cacheSynced(fakeCache(true))(req); err != nil {
|
||||
t.Errorf("synced: %v", err)
|
||||
}
|
||||
if err := cacheSynced(fakeCache(false))(req); err == nil {
|
||||
t.Error("unsynced caches reported ready")
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,126 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/imagepin"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"k8s.io/apimachinery/pkg/api/meta"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
)
|
||||
|
||||
// defaultRegistryURL is the [registry] url every install uses; deploy/bootstrap.sh
|
||||
// spells the same value as REGISTRY_URL.
|
||||
const defaultRegistryURL = "registry.felis.svc:5000"
|
||||
|
||||
// cmdPinImages pins every user server whose spec.image still names a mutable tag
|
||||
// in the platform registry to the digest that tag names now (internal/imagepin).
|
||||
// felis-api pins on create, so this covers the servers created before it did.
|
||||
//
|
||||
// deploy/bootstrap.sh runs it before it rebuilds the game images and pushes them
|
||||
// over the same tags: run after the push, it would pin those servers to the new
|
||||
// build, which is exactly the silent Minecraft upgrade pinning exists to stop.
|
||||
// It reaches the registry through the node's loopback hostPort, the same way the
|
||||
// installer pushes.
|
||||
func cmdPinImages(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("pin-images", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
namespace := fs.String("namespace", platform.DefaultMinecraftNamespace, "namespace the MinecraftServers live in")
|
||||
registry := fs.String("registry", defaultRegistryURL, "registry host[:port] the image refs spell")
|
||||
endpoint := fs.String("endpoint", "", "host[:port] to reach the registry at (default: 127.0.0.1 on the registry's port, its hostPort on this node)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return 0
|
||||
}
|
||||
return 2
|
||||
}
|
||||
if *endpoint == "" {
|
||||
*endpoint = loopbackEndpoint(*registry)
|
||||
}
|
||||
cl, err := buildSystemServerClient()
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis pin-images: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
|
||||
defer cancel()
|
||||
outcomes, err := pinUserServerImages(ctx, cl, *namespace, imagepin.Resolver{Registry: *registry, Endpoint: *endpoint})
|
||||
if meta.IsNoMatchError(err) {
|
||||
fmt.Fprintln(stdout, "felis pin-images: no MinecraftServer CRD yet, so no server to pin")
|
||||
return 0
|
||||
}
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis pin-images: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
if len(outcomes) == 0 {
|
||||
fmt.Fprintln(stdout, "felis pin-images: every user server already runs a pinned image")
|
||||
return 0
|
||||
}
|
||||
fmt.Fprintln(stdout, "felis pin-images: pinning user servers to the build their tag names now:")
|
||||
exit := 0
|
||||
for _, o := range outcomes {
|
||||
if o.err != nil {
|
||||
fmt.Fprintf(stdout, " - %s: ERROR %v\n", o.name, o.err)
|
||||
exit = 1
|
||||
continue
|
||||
}
|
||||
fmt.Fprintf(stdout, " - %s: %s\n", o.name, strings.Join(o.changes, ", "))
|
||||
}
|
||||
return exit
|
||||
}
|
||||
|
||||
// loopbackEndpoint is the registry's port on 127.0.0.1: the registry Deployment
|
||||
// binds it as a hostPort, and containerd's mirror and the installer's pushes use
|
||||
// the same address.
|
||||
func loopbackEndpoint(registry string) string {
|
||||
if i := strings.LastIndex(registry, ":"); i >= 0 {
|
||||
return "127.0.0.1" + registry[i:]
|
||||
}
|
||||
return "127.0.0.1"
|
||||
}
|
||||
|
||||
// pinUserServerImages patches spec.image of every user server whose image the
|
||||
// resolver covers and is not yet pinned. System servers are left on their tags:
|
||||
// the installer rebuilds and restarts them on purpose (restart_existing_system_servers).
|
||||
// A server that is already pinned, or runs an image from elsewhere, produces no
|
||||
// outcome, so a pinned fleet reports nothing. A running server restarts once as
|
||||
// the operator rolls its StatefulSet onto the pinned ref, which is the build it
|
||||
// already runs.
|
||||
func pinUserServerImages(ctx context.Context, cl client.Client, namespace string, r imagepin.Resolver) ([]systemServerOutcome, error) {
|
||||
var list v1alpha1.MinecraftServerList
|
||||
if err := cl.List(ctx, &list, client.InNamespace(namespace)); err != nil {
|
||||
return nil, fmt.Errorf("list servers: %w", err)
|
||||
}
|
||||
var out []systemServerOutcome
|
||||
for i := range list.Items {
|
||||
ms := &list.Items[i]
|
||||
if ms.Labels[v1alpha1.LabelSystemRole] != "" || imagepin.Pinned(ms.Spec.Image) || !r.Covers(ms.Spec.Image) {
|
||||
continue
|
||||
}
|
||||
pinned, err := r.Pin(ctx, ms.Spec.Image)
|
||||
if errors.Is(err, imagepin.ErrNotFound) {
|
||||
err = fmt.Errorf("%s is not in the registry, so there is no build to pin it to; left unpinned: %w", ms.Spec.Image, err)
|
||||
}
|
||||
if err != nil {
|
||||
out = append(out, systemServerOutcome{name: ms.Name, err: err})
|
||||
continue
|
||||
}
|
||||
patch := client.MergeFrom(ms.DeepCopy())
|
||||
ms.Spec.Image = pinned
|
||||
if err := cl.Patch(ctx, ms, patch); err != nil {
|
||||
out = append(out, systemServerOutcome{name: ms.Name, err: fmt.Errorf("patch %s: %w", ms.Name, err)})
|
||||
continue
|
||||
}
|
||||
out = append(out, systemServerOutcome{name: ms.Name, available: true, updated: true,
|
||||
changes: []string{"spec.image pinned to " + pinned}})
|
||||
}
|
||||
return out, nil
|
||||
}
|
||||
@@ -0,0 +1,112 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/imagepin"
|
||||
"felis.lolicon.best/internal/naming"
|
||||
"k8s.io/apimachinery/pkg/api/meta"
|
||||
"k8s.io/apimachinery/pkg/runtime/schema"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client/fake"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client/interceptor"
|
||||
)
|
||||
|
||||
const pinTestDigest = "sha256:2222222222222222222222222222222222222222222222222222222222222222"
|
||||
|
||||
// TestPinUserServerImages pins exactly the user servers still on a platform tag,
|
||||
// reports a tag the registry lost as an error without touching that server, and
|
||||
// has nothing left to do on a second pass.
|
||||
func TestPinUserServerImages(t *testing.T) {
|
||||
reg := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Path != "/v2/felis/paper/manifests/demo" {
|
||||
http.NotFound(w, r)
|
||||
return
|
||||
}
|
||||
w.Header().Set("Docker-Content-Digest", pinTestDigest)
|
||||
}))
|
||||
defer reg.Close()
|
||||
res := imagepin.Resolver{Registry: defaultRegistryURL, Endpoint: strings.TrimPrefix(reg.URL, "http://")}
|
||||
|
||||
paper := defaultRegistryURL + "/felis/paper:demo"
|
||||
mk := func(name, image, role string) *v1alpha1.MinecraftServer {
|
||||
ms := &v1alpha1.MinecraftServer{}
|
||||
ms.Name, ms.Namespace = name, "minecraft"
|
||||
ms.Spec.Image = image
|
||||
if role != "" {
|
||||
ms.Labels = map[string]string{v1alpha1.LabelSystemRole: role}
|
||||
}
|
||||
return ms
|
||||
}
|
||||
cl := fake.NewClientBuilder().WithScheme(newSystemServerScheme(t)).WithObjects(
|
||||
mk("legacy", paper, ""),
|
||||
mk("pinned", paper+"@sha256:"+strings.Repeat("3", 64), ""),
|
||||
mk("external", "docker.io/itzg/minecraft-server:java21", ""),
|
||||
mk("gone", defaultRegistryURL+"/felis/paper:old", ""),
|
||||
mk(naming.SystemLobbyServer, defaultRegistryURL+"/felis/felis-lobby:demo", naming.SystemLobbyServer),
|
||||
).Build()
|
||||
ctx := context.Background()
|
||||
|
||||
outcomes, err := pinUserServerImages(ctx, cl, "minecraft", res)
|
||||
if err != nil {
|
||||
t.Fatalf("pinUserServerImages: %v", err)
|
||||
}
|
||||
byName := map[string]systemServerOutcome{}
|
||||
for _, o := range outcomes {
|
||||
byName[o.name] = o
|
||||
}
|
||||
if len(outcomes) != 2 || byName["legacy"].err != nil || byName["gone"].err == nil {
|
||||
t.Fatalf("outcomes = %+v, want legacy pinned and gone reported", outcomes)
|
||||
}
|
||||
|
||||
want := map[string]string{
|
||||
"legacy": paper + "@" + pinTestDigest,
|
||||
"pinned": paper + "@sha256:" + strings.Repeat("3", 64),
|
||||
"external": "docker.io/itzg/minecraft-server:java21",
|
||||
"gone": defaultRegistryURL + "/felis/paper:old",
|
||||
naming.SystemLobbyServer: defaultRegistryURL + "/felis/felis-lobby:demo",
|
||||
}
|
||||
for name, image := range want {
|
||||
var ms v1alpha1.MinecraftServer
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: name}, &ms); err != nil {
|
||||
t.Fatalf("get %s: %v", name, err)
|
||||
}
|
||||
if ms.Spec.Image != image {
|
||||
t.Errorf("%s image = %q, want %q", name, ms.Spec.Image, image)
|
||||
}
|
||||
}
|
||||
|
||||
again, err := pinUserServerImages(ctx, cl, "minecraft", res)
|
||||
if err != nil || len(again) != 1 || again[0].name != "gone" {
|
||||
t.Fatalf("second pass = %+v, %v; want only the unresolvable server again", again, err)
|
||||
}
|
||||
}
|
||||
|
||||
// A fresh install has no CRD yet; the command must read that as nothing to pin.
|
||||
func TestPinUserServerImagesNoCRD(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(newSystemServerScheme(t)).WithInterceptorFuncs(interceptor.Funcs{
|
||||
List: func(context.Context, client.WithWatch, client.ObjectList, ...client.ListOption) error {
|
||||
return &meta.NoKindMatchError{GroupKind: schema.GroupKind{Group: "felis.lolicon.best", Kind: "MinecraftServer"}}
|
||||
},
|
||||
}).Build()
|
||||
_, err := pinUserServerImages(context.Background(), cl, "minecraft", imagepin.Resolver{Registry: defaultRegistryURL})
|
||||
if !meta.IsNoMatchError(err) {
|
||||
t.Fatalf("err = %v, want a NoMatch error the command can recognise", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestLoopbackEndpoint(t *testing.T) {
|
||||
for in, want := range map[string]string{
|
||||
"registry.felis.svc:5000": "127.0.0.1:5000",
|
||||
"registry.example": "127.0.0.1",
|
||||
} {
|
||||
if got := loopbackEndpoint(in); got != want {
|
||||
t.Errorf("loopbackEndpoint(%q) = %q, want %q", in, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
+176
-25
@@ -1,9 +1,13 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
@@ -12,8 +16,11 @@ import (
|
||||
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
||||
"felis.lolicon.best/internal/backup"
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/mail"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"felis.lolicon.best/internal/reaper"
|
||||
"felis.lolicon.best/internal/store"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
"k8s.io/apimachinery/pkg/runtime"
|
||||
utilruntime "k8s.io/apimachinery/pkg/util/runtime"
|
||||
clientgoscheme "k8s.io/client-go/kubernetes/scheme"
|
||||
@@ -30,7 +37,7 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("reaper", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
||||
worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>)")
|
||||
worldsRoot := fs.String("worlds-root", "/worlds", "mount root under which world PVCs are visible (tarLocal: <root>/<pvc>, else the stock local-path <root>/<pv-name>_<ns>_<pvc-name>)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
@@ -47,21 +54,8 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
|
||||
return 1
|
||||
}
|
||||
|
||||
archiver, err := buildArchiver(cfg, *worldsRoot)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
ctx := ctrl.SetupSignalHandler()
|
||||
|
||||
drv, err := store.Open(ctx, cfg.Database.URL)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer drv.Close()
|
||||
|
||||
scheme := runtime.NewScheme()
|
||||
utilruntime.Must(clientgoscheme.AddToScheme(scheme))
|
||||
utilruntime.Must(v1alpha1.AddToScheme(scheme))
|
||||
@@ -71,6 +65,19 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
|
||||
return 1
|
||||
}
|
||||
|
||||
archiver, err := buildArchiver(ctx, cfg, *worldsRoot, cl)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
|
||||
drv, err := store.Open(ctx, cfg.Database.URL)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis reaper: open database: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
defer drv.Close()
|
||||
|
||||
r := &reaper.Reaper{
|
||||
Cfg: rcfg,
|
||||
Store: reaper.NewPGStore(drv.DB()),
|
||||
@@ -78,19 +85,107 @@ func cmdReaper(args []string, stdout, stderr io.Writer) int {
|
||||
Archiver: archiver,
|
||||
}
|
||||
|
||||
// Pre-reap warnings go out by email when [smtp] is configured (the same
|
||||
// relay and password_ref convention felis-api uses); without it the channel
|
||||
// stays nil and the reaper logs each suppressed warning instead of stamping
|
||||
// it, so a later SMTP setup still gets to warn. The owner must have a
|
||||
// VERIFIED address — that flag is what proves the mailbox.
|
||||
if cfg.SMTP.Host != "" {
|
||||
passRef := cfg.SMTP.PasswordRef
|
||||
if passRef == "" {
|
||||
passRef = platform.SMTPPasswordEnv
|
||||
}
|
||||
password := os.Getenv(passRef)
|
||||
if cfg.SMTP.Username != "" && password == "" {
|
||||
fmt.Fprintf(stderr, "felis reaper: warning: [smtp] username is set but credentials env %s is empty — warning emails will fail AUTH\n", passRef)
|
||||
}
|
||||
db := drv.DB()
|
||||
r.Warner = &mailWarner{
|
||||
lookupEmail: func(ctx context.Context, ownerID string) (string, error) {
|
||||
var email string
|
||||
switch err := db.QueryRowContext(ctx,
|
||||
`SELECT email FROM users
|
||||
WHERE id = $1 AND email_verified = true AND COALESCE(email, '') <> ''`,
|
||||
ownerID).Scan(&email); {
|
||||
case errors.Is(err, sql.ErrNoRows):
|
||||
return "", fmt.Errorf("owner %s has no verified email", ownerID)
|
||||
case err != nil:
|
||||
return "", err
|
||||
}
|
||||
return email, nil
|
||||
},
|
||||
notifier: &mail.SMTP{
|
||||
Host: cfg.SMTP.Host,
|
||||
Port: cfg.SMTP.Port,
|
||||
From: cfg.SMTP.From,
|
||||
Username: cfg.SMTP.Username,
|
||||
Password: password,
|
||||
},
|
||||
}
|
||||
} else {
|
||||
fmt.Fprintln(stderr, "felis reaper: [smtp] not configured — pre-reap warnings are logged and NOT marked sent")
|
||||
}
|
||||
|
||||
sum, err := r.RunOnce(ctx)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis reaper: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis reaper: evaluated=%d reaped=%d warned=%d skipped=%d evicted=%d expired=%d\n",
|
||||
sum.Evaluated, sum.WorldsReaped, sum.Warned, sum.Skipped, sum.EvictedEarly, sum.BackupsExpired)
|
||||
return 0
|
||||
return reportReaperRun(sum, stdout, stderr)
|
||||
}
|
||||
|
||||
// reportReaperRun prints the run's tally and turns a run that left work undone
|
||||
// into exit 1, so the Job fails and the watchdog's job-failed check (and the
|
||||
// FelisWorldJobFailed rule) reach the operator: a world that cannot be archived
|
||||
// is kept, and without this nobody would learn that it is never reaped.
|
||||
func reportReaperRun(sum reaper.Summary, stdout, stderr io.Writer) int {
|
||||
fmt.Fprintf(stdout, "felis reaper: evaluated=%d reaped=%d awaiting_offsite=%d warned=%d skipped=%d store_full=%d evicted=%d expired=%d expire_failed=%d\n",
|
||||
sum.Evaluated, sum.WorldsReaped, sum.AwaitingOffsite, sum.Warned, sum.Skipped, sum.StoreFull,
|
||||
sum.EvictedEarly, sum.BackupsExpired, sum.ExpireFailed)
|
||||
if !sum.Failed() {
|
||||
return 0
|
||||
}
|
||||
fmt.Fprintf(stderr, "felis reaper: %d servers failed (%d kept because the backup store is full) and %d expired backups were not removed; the errors are above, and each is retried next run\n",
|
||||
sum.Skipped, sum.StoreFull, sum.ExpireFailed)
|
||||
return 1
|
||||
}
|
||||
|
||||
// mailWarner delivers a pre-reap notice to the owner's verified email — the
|
||||
// only channel this build can reach. Unowned owners and owners who never proved
|
||||
// a mailbox yield an error; the reaper retries such notices on its next run and
|
||||
// never lets them block the reap (red line ⑤).
|
||||
type mailWarner struct {
|
||||
lookupEmail func(ctx context.Context, ownerID string) (string, error)
|
||||
notifier noticeNotifier
|
||||
}
|
||||
|
||||
// noticeNotifier is the slice of mail.SMTP the warner needs (injected in tests).
|
||||
type noticeNotifier interface {
|
||||
SendNotice(ctx context.Context, email, subject, body string) error
|
||||
}
|
||||
|
||||
func (w *mailWarner) Warn(ctx context.Context, ownerID, server, remaining string) error {
|
||||
email, err := w.lookupEmail(ctx, ownerID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("resolve owner email: %w", err)
|
||||
}
|
||||
subject := fmt.Sprintf("Felis: 服务器 %s 将在 %s 后回收 · server reaped in %s", server, remaining, remaining)
|
||||
body := fmt.Sprintf(
|
||||
"Felis 世界回收提醒 / world-reaper notice\r\n"+
|
||||
"\r\n"+
|
||||
"服务器 / Server: %s\r\n"+
|
||||
"距回收 / Time left: %s\r\n"+
|
||||
"\r\n"+
|
||||
"闲置的服务器会先自动备份,再释放世界;有人加入游戏即可重置倒计时。\r\n"+
|
||||
"Idle servers are backed up and then released; any join resets the countdown.\r\n",
|
||||
server, remaining)
|
||||
return w.notifier.SendNotice(ctx, email, subject, body)
|
||||
}
|
||||
|
||||
// reaperConfig derives the reaper's retention windows from felis.toml. The 15d
|
||||
// idle deadline is fixed by §18; only the warning offsets, retention, and the
|
||||
// store soft-cap are configurable (§24).
|
||||
// idle deadline is fixed by §18; only the warning offsets, retention, the
|
||||
// store soft-cap and the on-demand backup bounds are configurable (§24). The
|
||||
// backup Job and felis-api read the manual_* bounds through it too.
|
||||
func reaperConfig(cfg *config.Config) (reaper.Config, error) {
|
||||
rc := reaper.DefaultConfig()
|
||||
if v := cfg.Archive.Retention; v != "" {
|
||||
@@ -118,26 +213,82 @@ func reaperConfig(cfg *config.Config) (reaper.Config, error) {
|
||||
}
|
||||
rc.MaxLocalBytes = b
|
||||
}
|
||||
if v := cfg.Archive.ManualRetention; v != "" {
|
||||
d, err := parseSpanDuration(v)
|
||||
if err != nil || d <= 0 {
|
||||
return rc, fmt.Errorf("[archive] manual_retention %q: want a positive span such as 30d", v)
|
||||
}
|
||||
rc.ManualRetention = d
|
||||
}
|
||||
switch n := cfg.Archive.ManualKeep; {
|
||||
case n < 0:
|
||||
return rc, fmt.Errorf("[archive] manual_keep %d: want 1 or more", n)
|
||||
case n > 0:
|
||||
rc.ManualKeep = n
|
||||
}
|
||||
if v := cfg.Archive.ManualCooldown; v != "" {
|
||||
d, err := parseSpanDuration(v)
|
||||
if err != nil || d < 0 {
|
||||
return rc, fmt.Errorf("[archive] manual_cooldown %q: want a span such as 10m (0s for none)", v)
|
||||
}
|
||||
rc.ManualCooldown = d
|
||||
}
|
||||
rc.RequireOffsite = cfg.Offsite.Enabled()
|
||||
return rc, nil
|
||||
}
|
||||
|
||||
// buildArchiver constructs the WorldArchiver. Only tarLocal is implemented in
|
||||
// this build; the resolver maps each world PVC to <worldsRoot>/<pvc>, the mount
|
||||
// convention the reaper Job is deployed with.
|
||||
func buildArchiver(cfg *config.Config, worldsRoot string) (backup.WorldArchiver, error) {
|
||||
// this build; the resolver maps each world PVC to its directory under worldsRoot
|
||||
// (resolveWorldDir).
|
||||
func buildArchiver(ctx context.Context, cfg *config.Config, worldsRoot string, cl client.Client) (backup.WorldArchiver, error) {
|
||||
switch cfg.Archive.Store {
|
||||
case "tarLocal":
|
||||
return &backup.TarLocal{
|
||||
BackupRoot: cfg.Archive.LocalPath,
|
||||
Resolve: func(pvc string) (string, error) {
|
||||
return filepath.Join(worldsRoot, pvc), nil
|
||||
},
|
||||
Resolve: resolveWorldDir(ctx, cl, cfg.K8s.Namespace, worldsRoot),
|
||||
}, nil
|
||||
default:
|
||||
return nil, fmt.Errorf("[archive] store %q is not implemented in this build (only tarLocal)", cfg.Archive.Store)
|
||||
}
|
||||
}
|
||||
|
||||
// resolveWorldDir maps a world PVC to its directory under worldsRoot, supporting
|
||||
// the two layouts a Felis host actually has:
|
||||
//
|
||||
// 1. <root>/<pvc> — the reaper's documented arrangement (worlds exposed by PVC
|
||||
// name, e.g. via mounting each volume or a crafted storage class).
|
||||
// 2. <root>/<pv-name>_<namespace>_<pvc-name> — what a stock k3s install gets:
|
||||
// local-path-provisioner stores every volume under its storage root as that
|
||||
// exact directory name. Without this arm, retention on a default install could
|
||||
// only ever fail to find a world (a no-op reaper, or worse an operator
|
||||
// arranging paths by hand).
|
||||
//
|
||||
// The second path is derived EXACTLY from the live PVC's spec.volumeName, never
|
||||
// from a glob: a leftover directory of an old, deleted PV must never be mistaken
|
||||
// for the world the PVC currently binds, because the reaper archives the resolved
|
||||
// directory and then deletes that PVC — archiving stale bytes and deleting the
|
||||
// real world would be data loss. When neither path exists the first is returned,
|
||||
// so the archive walk fails loudly against the documented path.
|
||||
func resolveWorldDir(ctx context.Context, cl client.Client, namespace, worldsRoot string) backup.PVCResolver {
|
||||
return func(pvc string) (string, error) {
|
||||
direct := filepath.Join(worldsRoot, pvc)
|
||||
if _, err := os.Stat(direct); err == nil {
|
||||
return direct, nil
|
||||
}
|
||||
var claim corev1.PersistentVolumeClaim
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: namespace, Name: pvc}, &claim); err != nil {
|
||||
return "", fmt.Errorf("resolve world PVC %s: %w", pvc, err)
|
||||
}
|
||||
if pv := claim.Spec.VolumeName; pv != "" {
|
||||
volDir := filepath.Join(worldsRoot, fmt.Sprintf("%s_%s_%s", pv, claim.Namespace, claim.Name))
|
||||
if _, err := os.Stat(volDir); err == nil {
|
||||
return volDir, nil
|
||||
}
|
||||
}
|
||||
return direct, nil
|
||||
}
|
||||
}
|
||||
|
||||
// parseSpanDuration parses the human spans used in felis.toml's [archive] table:
|
||||
// "3mo" (months≈30d), "15d" (days), or any time.ParseDuration unit ("12h").
|
||||
func parseSpanDuration(s string) (time.Duration, error) {
|
||||
|
||||
@@ -0,0 +1,186 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client/fake"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/reaper"
|
||||
)
|
||||
|
||||
// TestReportReaperRunFailsTheJob: a run that could not process a server, or
|
||||
// could not remove an expired backup, exits 1 so the Job shows as failed.
|
||||
func TestReportReaperRunFailsTheJob(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
sum reaper.Summary
|
||||
want int
|
||||
}{
|
||||
{"clean", reaper.Summary{Evaluated: 3, WorldsReaped: 1, AwaitingOffsite: 1}, 0},
|
||||
{"server failed", reaper.Summary{Evaluated: 3, Skipped: 1}, 1},
|
||||
{"store full", reaper.Summary{Evaluated: 3, Skipped: 1, StoreFull: 1}, 1},
|
||||
{"expiry failed", reaper.Summary{Evaluated: 3, ExpireFailed: 2}, 1},
|
||||
} {
|
||||
var out, errb bytes.Buffer
|
||||
if got := reportReaperRun(tc.sum, &out, &errb); got != tc.want {
|
||||
t.Errorf("%s: exit %d, want %d", tc.name, got, tc.want)
|
||||
}
|
||||
if !strings.Contains(out.String(), "skipped=") || !strings.Contains(out.String(), "expire_failed=") {
|
||||
t.Errorf("%s: summary line = %q", tc.name, out.String())
|
||||
}
|
||||
if (tc.want == 1) != (errb.Len() > 0) {
|
||||
t.Errorf("%s: stderr = %q", tc.name, errb.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestResolveWorldDir pins the two world layouts the reaper must find, and the
|
||||
// fail-closed miss. The stock local-path arm is derived from the live PVC's
|
||||
// volumeName — a name-based guess (glob) could tar a stale deleted PV's bytes and
|
||||
// then delete the current world, which is why it is read from the API instead.
|
||||
// TestReaperConfigManualKeys: the on-demand backup keys default to 30 days,
|
||||
// five per server and a ten-minute cooldown, accept overrides, and refuse
|
||||
// values that would keep nothing or throttle backwards.
|
||||
func TestReaperConfigManualKeys(t *testing.T) {
|
||||
rc, err := reaperConfig(&config.Config{})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if rc.ManualRetention != 30*reaper.Day || rc.ManualKeep != 5 || rc.ManualCooldown != 10*time.Minute {
|
||||
t.Fatalf("defaults = %v / %d / %v", rc.ManualRetention, rc.ManualKeep, rc.ManualCooldown)
|
||||
}
|
||||
rc, err = reaperConfig(&config.Config{Archive: config.ArchiveConfig{
|
||||
ManualRetention: "7d", ManualKeep: 2, ManualCooldown: "0s"}})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
if rc.ManualRetention != 7*reaper.Day || rc.ManualKeep != 2 || rc.ManualCooldown != 0 {
|
||||
t.Fatalf("overrides = %v / %d / %v", rc.ManualRetention, rc.ManualKeep, rc.ManualCooldown)
|
||||
}
|
||||
for _, bad := range []config.ArchiveConfig{
|
||||
{ManualRetention: "0d"},
|
||||
{ManualRetention: "soon"},
|
||||
{ManualKeep: -1},
|
||||
{ManualCooldown: "-5m"},
|
||||
{ManualCooldown: "often"},
|
||||
} {
|
||||
if _, err := reaperConfig(&config.Config{Archive: bad}); err == nil {
|
||||
t.Errorf("%+v was accepted", bad)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestResolveWorldDir(t *testing.T) {
|
||||
ctx := context.Background()
|
||||
root := t.TempDir()
|
||||
|
||||
// Arrange a world under the documented <root>/<pvc> layout.
|
||||
named := filepath.Join(root, "world-named-0")
|
||||
if err := os.MkdirAll(named, 0o750); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
// Arrange a second world the way k3s local-path stores it.
|
||||
pvDir := filepath.Join(root, "pvc-11111111-2222-3333-4444-555555555555_minecraft_world-live-0")
|
||||
if err := os.MkdirAll(pvDir, 0o750); err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
claim := &corev1.PersistentVolumeClaim{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: "world-live-0", Namespace: "minecraft"},
|
||||
Spec: corev1.PersistentVolumeClaimSpec{
|
||||
VolumeName: "pvc-11111111-2222-3333-4444-555555555555",
|
||||
},
|
||||
}
|
||||
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(claim).Build()
|
||||
resolve := resolveWorldDir(ctx, cl, "minecraft", root)
|
||||
|
||||
t.Run("documented name layout wins", func(t *testing.T) {
|
||||
got, err := resolve("world-named-0")
|
||||
if err != nil || got != named {
|
||||
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, named)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("stock local-path layout resolves exactly", func(t *testing.T) {
|
||||
got, err := resolve("world-live-0")
|
||||
if err != nil || got != pvDir {
|
||||
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, pvDir)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("neither layout present falls back to the documented path", func(t *testing.T) {
|
||||
// The claim exists but its directory does not: return the documented path so
|
||||
// the archive walk fails there, and the reaper preserves the world.
|
||||
missing := &corev1.PersistentVolumeClaim{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: "world-gone-0", Namespace: "minecraft"},
|
||||
Spec: corev1.PersistentVolumeClaimSpec{VolumeName: "pvc-99999999-0000-0000-0000-000000000000"},
|
||||
}
|
||||
cl := fake.NewClientBuilder().WithScheme(haltScheme(t)).WithObjects(missing).Build()
|
||||
got, err := resolveWorldDir(ctx, cl, "minecraft", root)("world-gone-0")
|
||||
if err != nil || got != filepath.Join(root, "world-gone-0") {
|
||||
t.Fatalf("resolve = (%q, %v), want (%q, nil)", got, err, filepath.Join(root, "world-gone-0"))
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("unknown pvc is an error, not a guess", func(t *testing.T) {
|
||||
_, err := resolve("world-unknown-0")
|
||||
if err == nil || !strings.Contains(err.Error(), "resolve world PVC world-unknown-0") {
|
||||
t.Fatalf("err = %v, want a resolve-world-PVC error", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// The pre-reap warner resolves the owner's VERIFIED email and hands the notice
|
||||
// to the mailer. Every failure (no verified address, relay refusal) returns an
|
||||
// error so the reaper retries on its next run instead of stamping a notice
|
||||
// nobody received.
|
||||
func TestMailWarner(t *testing.T) {
|
||||
lookup := func(email string, err error) func(context.Context, string) (string, error) {
|
||||
return func(context.Context, string) (string, error) { return email, err }
|
||||
}
|
||||
|
||||
n := &captureNotifier{}
|
||||
w := &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: n}
|
||||
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err != nil {
|
||||
t.Fatalf("Warn: %v", err)
|
||||
}
|
||||
if n.email != "[email protected]" || !strings.Contains(n.subject, "survival") || !strings.Contains(n.subject, "3d") {
|
||||
t.Fatalf("notice envelope = (%q, %q)", n.email, n.subject)
|
||||
}
|
||||
if !strings.Contains(n.body, "survival") || !strings.Contains(n.body, "3d") {
|
||||
t.Fatalf("body missing server/remaining:\n%s", n.body)
|
||||
}
|
||||
|
||||
w = &mailWarner{lookupEmail: lookup("", errors.New("owner u2 has no verified email")), notifier: n}
|
||||
if err := w.Warn(context.Background(), "u2", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "verified email") {
|
||||
t.Fatalf("unverified owner = %v, want the lookup error surfaced", err)
|
||||
}
|
||||
|
||||
w = &mailWarner{lookupEmail: lookup("[email protected]", nil), notifier: &captureNotifier{err: errors.New("relay down")}}
|
||||
if err := w.Warn(context.Background(), "u1", "survival", "3d"); err == nil || !strings.Contains(err.Error(), "relay down") {
|
||||
t.Fatalf("relay failure = %v, want it surfaced", err)
|
||||
}
|
||||
}
|
||||
|
||||
type captureNotifier struct {
|
||||
email, subject, body string
|
||||
err error
|
||||
}
|
||||
|
||||
func (n *captureNotifier) SendNotice(_ context.Context, email, subject, body string) error {
|
||||
if n.err != nil {
|
||||
return n.err
|
||||
}
|
||||
n.email, n.subject, n.body = email, subject, body
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,164 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"log/slog"
|
||||
"net"
|
||||
"net/http"
|
||||
"net/url"
|
||||
"os"
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/imagepush"
|
||||
"felis.lolicon.best/internal/registrygate"
|
||||
)
|
||||
|
||||
// cmdRegistryGate is the sidecar entrypoint in the registry pod: it owns the
|
||||
// registry port (and the loopback hostPort containerd pulls through), lets reads
|
||||
// through anonymously, and forwards writes to the loopback-only registry:2 only
|
||||
// for an authenticated principal allowed to write that repository. See
|
||||
// internal/registrygate for the policy.
|
||||
//
|
||||
// Tokens are files under --auth-dir, one per principal (platform, build, prune),
|
||||
// mounted from the registry-auth Secret. A missing file disables that principal:
|
||||
// writes fail closed while every pull keeps working, which is the right way round
|
||||
// for a registry the running workloads depend on.
|
||||
//
|
||||
// --maint-listen is the GC sidecar's read-only handshake (registrygate.MaintHandler).
|
||||
// It has no authentication, so it must name a loopback address; --maint-dir keeps
|
||||
// an open window across a gate restart.
|
||||
func cmdRegistryGate(args []string, _, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("registry-gate", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
listen := fs.String("listen", ":5000", "address the gate serves the registry API on")
|
||||
upstream := fs.String("upstream", "http://127.0.0.1:5001", "the loopback registry the gate forwards to")
|
||||
authDir := fs.String("auth-dir", "/etc/felis-registry-auth", "directory holding one token file per principal")
|
||||
maintListen := fs.String("maint-listen", "", "loopback address for the GC sidecar's read-only handshake (empty disables it)")
|
||||
maintDir := fs.String("maint-dir", "", "directory that keeps an open read-only window across a gate restart")
|
||||
quiet := fs.Duration("maint-quiet", registrygate.DefaultQuiet, "how long writes must be idle before a read-only window is granted")
|
||||
dataDir := fs.String("data-dir", "", "the registry's storage root, mounted read-only, for the manifest index (empty disables it)")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
if *maintListen != "" && !loopbackAddr(*maintListen) {
|
||||
fmt.Fprintf(stderr, "felis registry-gate: --maint-listen %q must be a loopback address: the handshake has no authentication\n", *maintListen)
|
||||
return 2
|
||||
}
|
||||
target, err := url.Parse(*upstream)
|
||||
if err != nil || target.Scheme == "" || target.Host == "" {
|
||||
fmt.Fprintf(stderr, "felis registry-gate: bad --upstream %q\n", *upstream)
|
||||
return 2
|
||||
}
|
||||
log := slog.New(slog.NewTextHandler(stderr, nil))
|
||||
tokens := map[string]string{}
|
||||
for _, p := range registrygate.Principals {
|
||||
b, err := os.ReadFile(filepath.Join(*authDir, p))
|
||||
tok := strings.TrimSpace(string(b))
|
||||
if err != nil || tok == "" {
|
||||
log.Warn("registry principal disabled: no token", "principal", p, "dir", *authDir)
|
||||
continue
|
||||
}
|
||||
tokens[p] = tok
|
||||
}
|
||||
|
||||
gate := registrygate.New(target, tokens, log)
|
||||
gate.SetQuiet(*quiet)
|
||||
gate.DataDir = *dataDir
|
||||
if *maintDir != "" {
|
||||
if err := gate.SetMaintenanceState(registrygate.MaintStatePath(*maintDir)); err != nil {
|
||||
// A corrupt file must not keep the registry from serving pulls.
|
||||
log.Warn("ignoring the saved read-only window", "err", err)
|
||||
}
|
||||
}
|
||||
srv := &http.Server{
|
||||
Addr: *listen,
|
||||
Handler: gate,
|
||||
ReadHeaderTimeout: 10 * time.Second,
|
||||
}
|
||||
var maint *http.Server
|
||||
if *maintListen != "" {
|
||||
maint = &http.Server{Addr: *maintListen, Handler: gate.MaintHandler(), ReadHeaderTimeout: 10 * time.Second}
|
||||
go func() {
|
||||
if err := maint.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
|
||||
log.Error("maintenance listener stopped; garbage collection cannot get a read-only window", "err", err)
|
||||
}
|
||||
}()
|
||||
}
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
go func() {
|
||||
<-ctx.Done()
|
||||
shutdown, cancel := context.WithTimeout(context.Background(), 10*time.Second)
|
||||
defer cancel()
|
||||
_ = srv.Shutdown(shutdown)
|
||||
if maint != nil {
|
||||
_ = maint.Shutdown(shutdown)
|
||||
}
|
||||
}()
|
||||
log.Info("registry gate listening", "addr", *listen, "upstream", target.String(), "principals", len(tokens))
|
||||
if err := srv.ListenAndServe(); err != nil && !errors.Is(err, http.ErrServerClosed) {
|
||||
fmt.Fprintf(stderr, "felis registry-gate: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
// loopbackAddr reports whether a host:port listen address binds loopback only.
|
||||
func loopbackAddr(addr string) bool {
|
||||
host, _, err := net.SplitHostPort(addr)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
if host == "localhost" {
|
||||
return true
|
||||
}
|
||||
ip := net.ParseIP(host)
|
||||
return ip != nil && ip.IsLoopback()
|
||||
}
|
||||
|
||||
// cmdPushImage is the build Job's publish step. It runs after Kaniko built the
|
||||
// image into a tarball (--no-push) and Trivy passed that tarball, and it is the
|
||||
// only container of the build pod that holds the registry credential — the one
|
||||
// executing the untrusted Dockerfile never sees it.
|
||||
func cmdPushImage(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("push-image", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
tarPath := fs.String("tar", "", "image tarball Kaniko wrote with --tar-path")
|
||||
ref := fs.String("ref", "", "host/repository:tag to publish it as")
|
||||
scheme := fs.String("scheme", "http", "registry scheme: http for the in-cluster registry, https otherwise")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
return 2
|
||||
}
|
||||
if *tarPath == "" || *ref == "" {
|
||||
fmt.Fprintln(stderr, "felis push-image: --tar and --ref are required")
|
||||
return 2
|
||||
}
|
||||
if *scheme != "http" && *scheme != "https" {
|
||||
fmt.Fprintf(stderr, "felis push-image: bad --scheme %q\n", *scheme)
|
||||
return 2
|
||||
}
|
||||
user := os.Getenv("FELIS_REGISTRY_USERNAME")
|
||||
pass := os.Getenv("FELIS_REGISTRY_PASSWORD")
|
||||
if user == "" || pass == "" {
|
||||
fmt.Fprintln(stderr, "felis push-image: FELIS_REGISTRY_USERNAME/FELIS_REGISTRY_PASSWORD are empty — the registry refuses anonymous writes")
|
||||
return 2
|
||||
}
|
||||
ctx, stop := signal.NotifyContext(context.Background(), os.Interrupt, syscall.SIGTERM)
|
||||
defer stop()
|
||||
p := &imagepush.Pusher{Scheme: *scheme, Username: user, Password: pass, Log: stderr}
|
||||
digest, err := p.Push(ctx, *tarPath, *ref)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis push-image: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintln(stdout, digest)
|
||||
return 0
|
||||
}
|
||||
+37
-17
@@ -11,7 +11,9 @@ Usage:
|
||||
felis <command> [flags]
|
||||
|
||||
Commands:
|
||||
migrate up Apply embedded database migrations under an advisory lock
|
||||
migrate up Apply embedded database migrations under an advisory lock (snapshots the database first)
|
||||
db Back up, verify, list and restore the control-plane database (backup|restore|verify|list|check)
|
||||
offsite Copy world archives and database bundles to an off-site bucket, and fetch them back (sync|status|list|fetch-db|fetch-worlds|keygen)
|
||||
operator Run the MinecraftServer controller-manager
|
||||
api Run the felis-api HTTP server
|
||||
nano Run the Felis-nano hasJoined multiplexer (multi-Yggdrasil, no control plane)
|
||||
@@ -19,9 +21,16 @@ Commands:
|
||||
restore Extract a world archive into a world volume (internal Job entrypoint)
|
||||
backup Archive a world into the backup store and record it (internal Job entrypoint)
|
||||
files List/read/write one file in a stopped server's world (internal Job entrypoint)
|
||||
egress-gate Hold a build pod until its egress NetworkPolicy is enforced (internal Job entrypoint)
|
||||
fetch-context Fetch and extract a submission's build context (internal Job entrypoint)
|
||||
push-image Push a scanned image tarball to the registry (internal Job entrypoint)
|
||||
mirror-build-tools Copy kaniko, trivy and Trivy's DBs into the registry (run by felis-build-tools.timer)
|
||||
registry-gate Authorize registry writes in front of registry:2 (internal sidecar entrypoint)
|
||||
manifests Render the control-plane RBAC + NetworkPolicy install bundle as YAML
|
||||
apply Create a MinecraftServer CRD (direct K8s write; use -f server.json)
|
||||
setup Run host bootstrap + first-run setup console (TUI; requires root/sudo)
|
||||
converge Fill in fields a newer desired spec added to already-installed system servers
|
||||
watchdog Check the platform once and mail the owners what has gone wrong (run by felis-watchdog.timer)
|
||||
version Print the build stamp of this binary
|
||||
update Report which platform components have updates available
|
||||
breakGlass Open the local break-glass emergency console (TUI; requires root/sudo)
|
||||
@@ -39,22 +48,33 @@ Run "felis <command> -h" for command-specific flags.
|
||||
// The help aliases are deliberately NOT entries: they print usage rather than run a
|
||||
// subcommand, and listing them would make the table disagree with the command list.
|
||||
var commands = map[string]func(args []string, stdout, stderr io.Writer) int{
|
||||
"migrate": cmdMigrate,
|
||||
"operator": cmdOperator,
|
||||
"api": cmdAPI,
|
||||
"nano": cmdNano,
|
||||
"reaper": cmdReaper,
|
||||
"restore": cmdRestore,
|
||||
"backup": cmdBackup,
|
||||
"files": cmdFiles,
|
||||
"manifests": cmdManifests,
|
||||
"apply": cmdApply,
|
||||
"setup": cmdSetup,
|
||||
"breakGlass": cmdBreakGlass,
|
||||
"bootstrap-assets": cmdBootstrapAssets,
|
||||
"init-forwarding": cmdInitForwarding,
|
||||
"version": cmdVersion,
|
||||
"update": cmdUpdate,
|
||||
"migrate": cmdMigrate,
|
||||
"db": cmdDB,
|
||||
"offsite": cmdOffsite,
|
||||
"operator": cmdOperator,
|
||||
"api": cmdAPI,
|
||||
"nano": cmdNano,
|
||||
"reaper": cmdReaper,
|
||||
"restore": cmdRestore,
|
||||
"backup": cmdBackup,
|
||||
"files": cmdFiles,
|
||||
"egress-gate": cmdEgressGate,
|
||||
"fetch-context": cmdFetchContext,
|
||||
"push-image": cmdPushImage,
|
||||
"mirror-build-tools": cmdMirrorBuildTools,
|
||||
"registry-gate": cmdRegistryGate,
|
||||
"manifests": cmdManifests,
|
||||
"apply": cmdApply,
|
||||
"setup": cmdSetup,
|
||||
"converge": cmdConverge,
|
||||
"breakGlass": cmdBreakGlass,
|
||||
"bootstrap-assets": cmdBootstrapAssets,
|
||||
"init-forwarding": cmdInitForwarding,
|
||||
"init-volume": cmdInitVolume,
|
||||
"pin-images": cmdPinImages,
|
||||
"version": cmdVersion,
|
||||
"update": cmdUpdate,
|
||||
"watchdog": cmdWatchdog,
|
||||
}
|
||||
|
||||
// run dispatches a subcommand. It is separate from main so the router is
|
||||
|
||||
@@ -38,9 +38,13 @@ func TestRunUnknownCommand(t *testing.T) {
|
||||
}
|
||||
|
||||
// undocumentedCommands are routable on purpose but kept out of the usage text: they
|
||||
// are called by deploy/bootstrap.sh, not by a human at a prompt. Listing them here is
|
||||
// what makes their absence from usage a deliberate decision rather than an oversight.
|
||||
var undocumentedCommands = map[string]bool{"bootstrap-assets": true, "init-forwarding": true}
|
||||
// are called by deploy/bootstrap.sh or the operator's initContainers, not by a human
|
||||
// at a prompt. Listing them here is what makes their absence from usage a deliberate
|
||||
// decision rather than an oversight.
|
||||
var undocumentedCommands = map[string]bool{
|
||||
"bootstrap-assets": true, "init-forwarding": true, "init-volume": true,
|
||||
"pin-images": true,
|
||||
}
|
||||
|
||||
// The usage text and the dispatch table must describe the same set of commands.
|
||||
//
|
||||
|
||||
+37
-7
@@ -225,16 +225,44 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
|
||||
// renamed it must replicate the Secret by hand.
|
||||
controlNS := platform.DefaultControlNamespace
|
||||
apiBaseURL := platform.InternalAPIBaseURL(controlNS)
|
||||
// Both Secrets must land in the minecraft namespace before the pods that mount
|
||||
// them are created: the service token (login authenticates to felis-api with it)
|
||||
// and the Velocity forwarding secret (every backend verifies the proxy's signed
|
||||
// handshake with it — without it the login gate would derive an OFFLINE UUID and
|
||||
// the Owner would bind the wrong Minecraft identity).
|
||||
// These Secrets must land in the minecraft namespace before the pods that
|
||||
// mount them are created: the service token (login authenticates to felis-api
|
||||
// with it), the Velocity forwarding secret (every backend verifies the proxy's
|
||||
// signed handshake with it — without it the login gate would derive an OFFLINE
|
||||
// UUID and the Owner would bind the wrong Minecraft identity), and felis-config
|
||||
// (the on-demand BACKUP Job runs in the minecraft namespace and mounts it to
|
||||
// self-record its world_backups row; without the replica the Job's volume
|
||||
// mount fails and every backup request strands in the cluster).
|
||||
// An empty build_namespace means the build system's compiled-in default; the
|
||||
// replica must target the namespace the Jobs actually run in.
|
||||
buildNS := cfg.Registry.BuildNamespace
|
||||
if buildNS == "" {
|
||||
buildNS = platform.DefaultBuildNamespace
|
||||
}
|
||||
secretOutcomes := []systemServerOutcome{
|
||||
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
|
||||
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token"),
|
||||
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false),
|
||||
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
|
||||
naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret"),
|
||||
naming.ForwardingSecretName, naming.ForwardingSecretKey, "forwarding-secret", "minecraft ns", false),
|
||||
// refresh=true: felis-config is the rendered config, not a credential. The
|
||||
// backup/restore/fileedit Jobs and the reaper mount this copy, so a re-run
|
||||
// must update it when the control plane's render has moved on (a stale copy
|
||||
// e.g. keeps an old database URL after a credential rotation).
|
||||
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
|
||||
"felis-config", "felis.toml", "config", "minecraft ns", true),
|
||||
// The reaper's pre-reap warning emails authenticate with the same relay
|
||||
// password felis-api uses; the reaper pod runs in the minecraft namespace,
|
||||
// where a secretKeyRef resolves only against a local mirror. Skipped while
|
||||
// the relay is not configured yet — the "configure email" screen refreshes
|
||||
// both mirrors when it applies.
|
||||
ensureSecretReplica(ctx, cl, controlNS, cfg.K8s.Namespace,
|
||||
"felis-smtp", "password", "smtp", "minecraft ns", false),
|
||||
// The build namespace needs the same token: the build Job's fetch
|
||||
// initContainer reads the submission context from the internal face. Best
|
||||
// effort — a deployment that only installs the control plane simply never
|
||||
// builds a user submission.
|
||||
ensureSecretReplica(ctx, cl, controlNS, buildNS,
|
||||
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "felis-build ns", false),
|
||||
}
|
||||
outcomes := ensureSystemServers(ctx, cl, cfg.K8s.Namespace, cfg.Velocity.LoginImage, cfg.Velocity.LobbyImage, apiBaseURL, cfg.Server.RootDomain, defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname))
|
||||
outcomes = append(secretOutcomes, outcomes...)
|
||||
@@ -245,6 +273,8 @@ func provisionSystemServers(ctx context.Context, cfg *config.Config, out io.Writ
|
||||
fmt.Fprintf(out, " - %s: ERROR %v\n", o.name, o.err)
|
||||
case o.created:
|
||||
fmt.Fprintf(out, " - %s: created (DesiredState=Running)\n", o.name)
|
||||
case o.updated:
|
||||
fmt.Fprintf(out, " - %s: refreshed from the control namespace\n", o.name)
|
||||
default:
|
||||
fmt.Fprintf(out, " - %s: skipped (%s)\n", o.name, o.skipped)
|
||||
}
|
||||
|
||||
+194
-35
@@ -1,6 +1,7 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"fmt"
|
||||
"time"
|
||||
@@ -95,7 +96,7 @@ const felisLimboHealthPort int32 = 8080
|
||||
// fail-safes to readiness-only, so a login pod that has the URL/domain but not yet
|
||||
// the token is safe (it simply does not authenticate) rather than broken.
|
||||
const (
|
||||
envAPIBaseURL = "FELIS_API_BASE_URL"
|
||||
envAPIBaseURL = naming.EnvAPIBaseURL
|
||||
envRootDomain = "FELIS_ROOT_DOMAIN"
|
||||
envPanelHostname = "FELIS_PANEL_HOSTNAME"
|
||||
envLobbyServer = "FELIS_LOBBY_SERVER"
|
||||
@@ -255,10 +256,31 @@ func buildSystemServerClient() (client.Client, error) {
|
||||
// setup can report it without the provisioner deciding on the output format.
|
||||
type systemServerOutcome struct {
|
||||
name string
|
||||
created bool // true = we created it this run
|
||||
available bool // true = the required object now exists
|
||||
skipped string // non-empty = why it was skipped (image unset / already exists)
|
||||
err error // non-nil = create failed
|
||||
created bool // true = we created it this run
|
||||
updated bool // true = we refreshed an existing replica from the source
|
||||
available bool // true = the required object now exists
|
||||
skipped string // non-empty = why it was skipped (image unset / already exists)
|
||||
err error // non-nil = create failed
|
||||
changes []string // converge only: the fields this pass filled
|
||||
}
|
||||
|
||||
// systemServerPlan is one system service in the provisioner's table: its name,
|
||||
// the image config gives it, and the pure builder for its desired CR.
|
||||
type systemServerPlan struct {
|
||||
name string
|
||||
image string
|
||||
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
|
||||
}
|
||||
|
||||
// systemServerPlans is the single description of the login+lobby pair, shared by
|
||||
// ensureSystemServers (create-if-absent) and convergeSystemServers (field fill).
|
||||
func systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerPlan {
|
||||
return []systemServerPlan{
|
||||
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
|
||||
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
|
||||
}},
|
||||
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
|
||||
}
|
||||
}
|
||||
|
||||
// ensureSystemServers idempotently creates the login and lobby system services.
|
||||
@@ -269,17 +291,7 @@ type systemServerOutcome struct {
|
||||
// K8s client and namespace; this function performs no signal-handler or client
|
||||
// setup of its own.
|
||||
func ensureSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
|
||||
type plan struct {
|
||||
name string
|
||||
image string
|
||||
build func(image, namespace string) (*v1alpha1.MinecraftServer, error)
|
||||
}
|
||||
plans := []plan{
|
||||
{name: naming.SystemLoginServer, image: loginImage, build: func(image, ns string) (*v1alpha1.MinecraftServer, error) {
|
||||
return loginSystemServer(image, ns, apiBaseURL, rootDomain, panelHostname)
|
||||
}},
|
||||
{name: naming.SystemLobbyServer, image: lobbyImage, build: lobbySystemServer},
|
||||
}
|
||||
plans := systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname)
|
||||
|
||||
outcomes := make([]systemServerOutcome, 0, len(plans))
|
||||
for _, p := range plans {
|
||||
@@ -379,12 +391,7 @@ var derivedSystemEnv = map[string]bool{
|
||||
// deliberate removal is indistinguishable from drift and re-adding it would fight the
|
||||
// operator every run.
|
||||
func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired *v1alpha1.MinecraftServer) (bool, error) {
|
||||
want := make(map[string]string, len(derivedSystemEnv))
|
||||
for _, e := range desired.Spec.Env {
|
||||
if derivedSystemEnv[e.Name] {
|
||||
want[e.Name] = e.Value
|
||||
}
|
||||
}
|
||||
want := derivedEnvWanted(desired)
|
||||
|
||||
changed := false
|
||||
for i, e := range existing.Spec.Env {
|
||||
@@ -402,6 +409,116 @@ func refreshDerivedEnv(ctx context.Context, cl client.Client, existing, desired
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// derivedEnvWanted maps the derived env keys of desired onto their values.
|
||||
func derivedEnvWanted(desired *v1alpha1.MinecraftServer) map[string]string {
|
||||
want := make(map[string]string, len(derivedSystemEnv))
|
||||
for _, e := range desired.Spec.Env {
|
||||
if derivedSystemEnv[e.Name] {
|
||||
want[e.Name] = e.Value
|
||||
}
|
||||
}
|
||||
return want
|
||||
}
|
||||
|
||||
// convergeSystemServers is the explicit convergence pass over already-installed
|
||||
// system servers (#1). ensureSystemServers is create-if-absent by design — an
|
||||
// existing CR is left alone so a re-run cannot clobber an operator's edits — and
|
||||
// that leaves no path for a field the DESIRED spec gained after the install:
|
||||
// spec.rcon (the lobby's write channel), spec.startup.healthHTTPPort (the login
|
||||
// gate's readiness probe), or a config-derived env key that did not exist yet.
|
||||
// Such fields sit at their zero value forever while re-running setup reports
|
||||
// success, which is exactly the reported "configuration updates never reach an
|
||||
// installed deployment" symptom.
|
||||
//
|
||||
// This pass fills exactly those zero-value fields and the config-derived env keys,
|
||||
// and nothing else: a field already holding a non-zero value is the operator's and
|
||||
// is never overwritten. It is an explicit command rather than an implicit step of
|
||||
// setup because some fills need an ordering only the operator knows — enabling
|
||||
// RCON or the HTTP readiness gate on a server whose image predates the listener
|
||||
// would hold that server in Starting until it was marked Failed. Rebuild (or
|
||||
// upgrade) the images first, then run this.
|
||||
func convergeSystemServers(ctx context.Context, cl client.Client, namespace, loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname string) []systemServerOutcome {
|
||||
outcomes := make([]systemServerOutcome, 0, 2)
|
||||
for _, p := range systemServerPlans(loginImage, lobbyImage, apiBaseURL, rootDomain, panelHostname) {
|
||||
if p.image == "" {
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, skipped: "image not configured"})
|
||||
continue
|
||||
}
|
||||
desired, err := p.build(p.image, namespace)
|
||||
if err != nil {
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
|
||||
continue
|
||||
}
|
||||
|
||||
var existing v1alpha1.MinecraftServer
|
||||
switch err := cl.Get(ctx, client.ObjectKeyFromObject(desired), &existing); {
|
||||
case apierrors.IsNotFound(err):
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name,
|
||||
skipped: "not present — run `sudo felis setup` first"})
|
||||
continue
|
||||
case err != nil:
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: err})
|
||||
continue
|
||||
}
|
||||
if existing.Labels[v1alpha1.LabelSystemRole] != p.name {
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf(
|
||||
"existing MinecraftServer %s/%s is not marked as the Felis %q system role; refusing to converge it",
|
||||
namespace, p.name, p.name,
|
||||
)})
|
||||
continue
|
||||
}
|
||||
|
||||
var changes []string
|
||||
if existing.Spec.Rcon == (v1alpha1.RconSpec{}) && desired.Spec.Rcon != (v1alpha1.RconSpec{}) {
|
||||
existing.Spec.Rcon = desired.Spec.Rcon
|
||||
changes = append(changes, "spec.rcon")
|
||||
}
|
||||
if existing.Spec.Startup.HealthHTTPPort == 0 && desired.Spec.Startup.HealthHTTPPort != 0 {
|
||||
existing.Spec.Startup.HealthHTTPPort = desired.Spec.Startup.HealthHTTPPort
|
||||
changes = append(changes, "spec.startup.healthHTTPPort")
|
||||
}
|
||||
changes = append(changes, convergeDerivedEnv(&existing, desired)...)
|
||||
|
||||
if len(changes) == 0 {
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, skipped: "already converged"})
|
||||
continue
|
||||
}
|
||||
if err := cl.Update(ctx, &existing); err != nil {
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, err: fmt.Errorf("converge %s: %w", p.name, err)})
|
||||
continue
|
||||
}
|
||||
outcomes = append(outcomes, systemServerOutcome{name: p.name, available: true, updated: true, changes: changes})
|
||||
}
|
||||
return outcomes
|
||||
}
|
||||
|
||||
// convergeDerivedEnv makes the config-derived env match the desired values: a key
|
||||
// whose value drifted is overwritten, and a key missing entirely is added. This is
|
||||
// the wider half of the same explicit pass — refreshDerivedEnv's present-only loop
|
||||
// can never introduce a NEW key, which is how a derived key added after an install
|
||||
// never reached it at all.
|
||||
func convergeDerivedEnv(existing, desired *v1alpha1.MinecraftServer) []string {
|
||||
want := derivedEnvWanted(desired)
|
||||
var changes []string
|
||||
present := make(map[string]bool, len(existing.Spec.Env))
|
||||
for i := range existing.Spec.Env {
|
||||
e := &existing.Spec.Env[i]
|
||||
present[e.Name] = true
|
||||
if v, ok := want[e.Name]; ok && v != e.Value {
|
||||
e.Value = v
|
||||
changes = append(changes, "env "+e.Name)
|
||||
}
|
||||
}
|
||||
for _, e := range desired.Spec.Env {
|
||||
if !derivedSystemEnv[e.Name] || present[e.Name] {
|
||||
continue
|
||||
}
|
||||
existing.Spec.Env = append(existing.Spec.Env, e)
|
||||
changes = append(changes, "env "+e.Name)
|
||||
}
|
||||
return changes
|
||||
}
|
||||
|
||||
// The login gate is a hard prerequisite of the Owner bind, so setup waits for it
|
||||
// rather than racing it. The ceiling covers a cold image pull on a fresh node;
|
||||
// the poll is fast enough that a warm start feels immediate.
|
||||
@@ -471,26 +588,35 @@ func phaseOrPending(p v1alpha1.Phase) string {
|
||||
return string(p)
|
||||
}
|
||||
|
||||
// ensureSecretReplica copies one Secret from the control namespace into the minecraft
|
||||
// namespace so a backend pod can mount it via secretKeyRef. A secretKeyRef is
|
||||
// namespace-local, but the backends run in the minecraft namespace while the sources
|
||||
// of truth live beside the control plane — so without this replica the operator's
|
||||
// injected secretKeyRef would dangle and wedge the pod in CreateContainerConfigError.
|
||||
// ensureSecretReplica copies one Secret from the control namespace into a workload
|
||||
// namespace (minecraft — or the build namespace, whose fetch initContainer reads the
|
||||
// context from the felis-api internal face with the same token) so a pod can mount it
|
||||
// via secretKeyRef. A secretKeyRef is namespace-local, but those workloads do not run
|
||||
// beside the control plane — so without this replica the secretKeyRef would dangle and
|
||||
// wedge the pod in CreateContainerConfigError.
|
||||
//
|
||||
// Two Secrets need it, for different reasons: the service token (login only — it
|
||||
// authenticates the limbo plugin to the felis-api internal face) and the Velocity
|
||||
// modern-forwarding secret (every backend — it is how a backend knows a login really
|
||||
// came from the proxy, and so that the player's UUID is Mojang-verified rather than
|
||||
// offline-derived).
|
||||
// Three Secrets need it, for different reasons: the service token (the login limbo and
|
||||
// the build Pod's context fetch — both authenticate to the felis-api internal face),
|
||||
// the Velocity modern-forwarding secret (every backend — it is how a backend knows
|
||||
// a login really came from the proxy, and so that the player's UUID is Mojang-verified
|
||||
// rather than offline-derived), and the SMTP relay password (the reaper's pre-reap
|
||||
// warning emails; the felis-config mirror is what carries [smtp] into its pod).
|
||||
//
|
||||
// It is create-if-absent: an existing replica is left untouched so a hand-rotated
|
||||
// value in the minecraft namespace is never clobbered (to rotate, delete the replica
|
||||
// value in the workload namespace is never clobbered (to rotate, delete the replica
|
||||
// and re-run setup). Best-effort like the rest of the provisioner: a missing source or
|
||||
// a create failure degrades to a reported outcome, never a hard setup failure. It
|
||||
// copies only Type and Data — never labels/annotations/ownerRefs — so the replica
|
||||
// carries no accidental GC owner or managed-by lineage.
|
||||
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label string) systemServerOutcome {
|
||||
name := label + " (minecraft ns)"
|
||||
//
|
||||
// refreshExisting switches the felis-config mirror to refresh-in-place: that Secret is
|
||||
// a rendered config, never a hand-rotated credential, and the workload Jobs that mount
|
||||
// it (backup/restore/fileedit) plus the reaper silently misbehave on a stale copy —
|
||||
// e.g. after a database credential rotation the control plane moves on while every
|
||||
// backup Job keeps failing auth. Credential Secrets keep the never-overwrite rule so a
|
||||
// rotated value survives; to rotate those, delete the replica and re-run setup.
|
||||
func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace, minecraftNamespace, secretName, secretKey, label, where string, refreshExisting bool) systemServerOutcome {
|
||||
name := label + " (" + where + ")"
|
||||
validate := func(secret *corev1.Secret, location, skipped string) systemServerOutcome {
|
||||
if len(secret.Data[secretKey]) == 0 {
|
||||
return systemServerOutcome{name: name, skipped: fmt.Sprintf(
|
||||
@@ -498,6 +624,33 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
|
||||
}
|
||||
return systemServerOutcome{name: name, available: true, skipped: skipped}
|
||||
}
|
||||
// refreshFromControl updates an existing replica from the control-namespace source
|
||||
// when the rendered key differs. Only the felis-config mirror opts in.
|
||||
refreshFromControl := func(existing *corev1.Secret) systemServerOutcome {
|
||||
var src corev1.Secret
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: controlNamespace, Name: secretName}, &src); err != nil {
|
||||
if apierrors.IsNotFound(err) {
|
||||
return systemServerOutcome{name: name, skipped: fmt.Sprintf(
|
||||
"source Secret %s/%s not found — provision it (deploy/bootstrap.sh), then re-run setup",
|
||||
controlNamespace, secretName)}
|
||||
}
|
||||
return systemServerOutcome{name: name, err: err}
|
||||
}
|
||||
if out := validate(&src, controlNamespace, ""); !out.available {
|
||||
return out
|
||||
}
|
||||
if bytes.Equal(existing.Data[secretKey], src.Data[secretKey]) {
|
||||
return validate(existing, minecraftNamespace, "already current")
|
||||
}
|
||||
if existing.Data == nil {
|
||||
existing.Data = map[string][]byte{}
|
||||
}
|
||||
existing.Data[secretKey] = src.Data[secretKey]
|
||||
if err := cl.Update(ctx, existing); err != nil {
|
||||
return systemServerOutcome{name: name, err: err}
|
||||
}
|
||||
return systemServerOutcome{name: name, updated: true, available: true}
|
||||
}
|
||||
if controlNamespace == minecraftNamespace {
|
||||
// Same namespace needs no replica, but the source still has to exist.
|
||||
var existing corev1.Secret
|
||||
@@ -516,6 +669,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
|
||||
var existing corev1.Secret
|
||||
getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing)
|
||||
if getErr == nil {
|
||||
if refreshExisting {
|
||||
return refreshFromControl(&existing)
|
||||
}
|
||||
return validate(&existing, minecraftNamespace, "already exists")
|
||||
}
|
||||
if !apierrors.IsNotFound(getErr) {
|
||||
@@ -544,6 +700,9 @@ func ensureSecretReplica(ctx context.Context, cl client.Client, controlNamespace
|
||||
if getErr := cl.Get(ctx, client.ObjectKey{Namespace: minecraftNamespace, Name: secretName}, &existing); getErr != nil {
|
||||
return systemServerOutcome{name: name, err: getErr}
|
||||
}
|
||||
if refreshExisting {
|
||||
return refreshFromControl(&existing)
|
||||
}
|
||||
return validate(&existing, minecraftNamespace, "already exists")
|
||||
}
|
||||
return systemServerOutcome{name: name, err: err}
|
||||
|
||||
@@ -156,7 +156,7 @@ func TestEnsureSecretReplica(t *testing.T) {
|
||||
}
|
||||
replicate := func(cl client.Client, controlNS, mcNS string) systemServerOutcome {
|
||||
return ensureSecretReplica(ctx, cl, controlNS, mcNS,
|
||||
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token")
|
||||
naming.ServiceTokenSecretName, naming.ServiceTokenSecretKey, "service-token", "minecraft ns", false)
|
||||
}
|
||||
|
||||
t.Run("replicates when absent", func(t *testing.T) {
|
||||
@@ -251,6 +251,87 @@ func TestEnsureSecretReplica(t *testing.T) {
|
||||
})
|
||||
}
|
||||
|
||||
// The felis-config mirror is the one replica that must refresh: it is a rendered
|
||||
// config, and a stale workload-side copy (backup/restore/fileedit Jobs, the reaper)
|
||||
// misbehaves silently — a rotated database credential keeps the control plane moving
|
||||
// while every backup Job keeps failing auth. Credential Secrets keep create-if-absent.
|
||||
func TestEnsureSecretReplicaRefresh(t *testing.T) {
|
||||
scheme := newSystemServerScheme(t)
|
||||
ctx := context.Background()
|
||||
configSecret := func(ns, body string) *corev1.Secret {
|
||||
return &corev1.Secret{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: ns},
|
||||
Type: corev1.SecretTypeOpaque,
|
||||
Data: map[string][]byte{"felis.toml": []byte(body)},
|
||||
}
|
||||
}
|
||||
refresh := func(cl client.Client) systemServerOutcome {
|
||||
return ensureSecretReplica(ctx, cl, "felis", "minecraft",
|
||||
"felis-config", "felis.toml", "config", "minecraft ns", true)
|
||||
}
|
||||
replicaBody := func(t *testing.T, cl client.Client) string {
|
||||
t.Helper()
|
||||
var got corev1.Secret
|
||||
if err := cl.Get(ctx, client.ObjectKey{Namespace: "minecraft", Name: "felis-config"}, &got); err != nil {
|
||||
t.Fatalf("get replica: %v", err)
|
||||
}
|
||||
return string(got.Data["felis.toml"])
|
||||
}
|
||||
|
||||
t.Run("refreshes a stale config replica", func(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
|
||||
configSecret("felis", "current"),
|
||||
configSecret("minecraft", "stale"),
|
||||
).Build()
|
||||
out := refresh(cl)
|
||||
if out.err != nil || !out.updated || !out.available {
|
||||
t.Fatalf("outcome = %+v, want refreshed", out)
|
||||
}
|
||||
if got := replicaBody(t, cl); got != "current" {
|
||||
t.Errorf("replica = %q, want current", got)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("leaves a current config replica alone", func(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
|
||||
configSecret("felis", "same"),
|
||||
configSecret("minecraft", "same"),
|
||||
).Build()
|
||||
out := refresh(cl)
|
||||
if out.err != nil || out.updated || !out.available || out.skipped != "already current" {
|
||||
t.Fatalf("outcome = %+v, want already current", out)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("fills an empty-key replica", func(t *testing.T) {
|
||||
empty := &corev1.Secret{
|
||||
ObjectMeta: metav1.ObjectMeta{Name: "felis-config", Namespace: "minecraft"},
|
||||
Data: map[string][]byte{},
|
||||
}
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
|
||||
configSecret("felis", "current"), empty).Build()
|
||||
out := refresh(cl)
|
||||
if out.err != nil || !out.updated {
|
||||
t.Fatalf("outcome = %+v, want refreshed", out)
|
||||
}
|
||||
if got := replicaBody(t, cl); got != "current" {
|
||||
t.Errorf("replica = %q, want current", got)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("missing source degrades to a skip", func(t *testing.T) {
|
||||
cl := fake.NewClientBuilder().WithScheme(scheme).WithObjects(
|
||||
configSecret("minecraft", "stale")).Build()
|
||||
out := refresh(cl)
|
||||
if out.err != nil || out.updated || out.available || out.skipped == "" {
|
||||
t.Fatalf("outcome = %+v, want skipped (source missing)", out)
|
||||
}
|
||||
if got := replicaBody(t, cl); got != "stale" {
|
||||
t.Errorf("replica = %q, want untouched stale", got)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
func TestRequiredProvisioningError(t *testing.T) {
|
||||
ready := []systemServerOutcome{
|
||||
{name: "service-token (minecraft ns)", available: true},
|
||||
@@ -497,7 +578,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
|
||||
// human would have added and a memory bump off the built-in default.
|
||||
stale := func() *v1alpha1.MinecraftServer {
|
||||
ms, err := loginSystemServer("felis-limbo:demo", "minecraft",
|
||||
"http://old.internal:8081", "159.223.32.51.nip.io", "console.159.223.32.51.nip.io")
|
||||
"http://old.internal:8081", "203.0.113.10.nip.io", "console.203.0.113.10.nip.io")
|
||||
if err != nil {
|
||||
t.Fatalf("build stale login server: %v", err)
|
||||
}
|
||||
@@ -509,7 +590,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
|
||||
run := func(cl client.Client) []systemServerOutcome {
|
||||
return ensureSystemServers(ctx, cl, "minecraft", "felis-limbo:demo", "felis-lobby:demo",
|
||||
"http://felis-api-internal.felis.svc.cluster.local:8081",
|
||||
"mc.flyemoji.network", "console.mc.flyemoji.network")
|
||||
"mc.example.net", "console.mc.example.net")
|
||||
}
|
||||
|
||||
envOf := func(t *testing.T, cl client.Client) map[string]string {
|
||||
@@ -537,11 +618,11 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
|
||||
}
|
||||
}
|
||||
env := envOf(t, cl)
|
||||
if env[envPanelHostname] != "console.mc.flyemoji.network" {
|
||||
if env[envPanelHostname] != "console.mc.example.net" {
|
||||
t.Errorf("%s = %q — players are still being sent to the old console",
|
||||
envPanelHostname, env[envPanelHostname])
|
||||
}
|
||||
if env[envRootDomain] != "mc.flyemoji.network" {
|
||||
if env[envRootDomain] != "mc.example.net" {
|
||||
t.Errorf("%s = %q, want the new root domain", envRootDomain, env[envRootDomain])
|
||||
}
|
||||
})
|
||||
@@ -568,7 +649,7 @@ func TestEnsureSystemServersRefreshesDerivedEnv(t *testing.T) {
|
||||
t.Run("reports no refresh when config already matches", func(t *testing.T) {
|
||||
fresh, err := loginSystemServer("felis-limbo:demo", "minecraft",
|
||||
"http://felis-api-internal.felis.svc.cluster.local:8081",
|
||||
"mc.flyemoji.network", "console.mc.flyemoji.network")
|
||||
"mc.example.net", "console.mc.example.net")
|
||||
if err != nil {
|
||||
t.Fatalf("build fresh login server: %v", err)
|
||||
}
|
||||
|
||||
@@ -86,7 +86,7 @@ func (m *backupModel) loadCmd() tea.Cmd {
|
||||
if err != nil {
|
||||
return backupListMsg{err: fmt.Errorf("list servers: %w", err)}
|
||||
}
|
||||
return backupListMsg{cl: cl, servers: servers}
|
||||
return backupListMsg{cl: cl, servers: backupPickable(servers)}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -437,6 +437,10 @@ func (m *edgeModel) errorView() string {
|
||||
if m.lastErr != nil {
|
||||
b.WriteString(tuiHint.Render(m.lastErr.Error()) + "\n")
|
||||
}
|
||||
// Nothing done above is rolled back, and nothing needs to be: every step finds what an
|
||||
// earlier attempt created (the tunnel, its DNS route, the Access app and policy) and
|
||||
// carries on from it.
|
||||
b.WriteString("\n" + tuiHint.Render("Retrying is safe: it reuses the tunnel, DNS record and Access app created so far instead of making duplicates.") + "\n")
|
||||
b.WriteString("\n" + tuiAction("enter", "retry", "esc", "edit"))
|
||||
return b.String()
|
||||
}
|
||||
|
||||
@@ -32,7 +32,9 @@ func applyCloudflareEdge(ctx context.Context, result *cfsetup.Result, panelHost,
|
||||
if adminHost == "" {
|
||||
return fmt.Errorf("admin hostname is required")
|
||||
}
|
||||
if err := writeConnectionConfig(panelHost, adminHost, result.AccessAud); err != nil {
|
||||
// cloudflared is the only way in once the NodePort is fenced, so the
|
||||
// visitor address it writes can key the sign-in rate limit.
|
||||
if err := writeConnectionConfig(panelHost, adminHost, result.AccessAud, "CF-Connecting-IP"); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := applyFelisConfigSecret(ctx); err != nil {
|
||||
@@ -78,7 +80,8 @@ func applyReverseProxy(ctx context.Context, panelHost, adminHost string) error {
|
||||
if adminHost == "" {
|
||||
return fmt.Errorf("admin hostname is required")
|
||||
}
|
||||
if err := writeConnectionConfig(panelHost, adminHost, ""); err != nil {
|
||||
// Caddy, nginx and Traefik all append the peer they saw to X-Forwarded-For.
|
||||
if err := writeConnectionConfig(panelHost, adminHost, "", "X-Forwarded-For"); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := applyFelisConfigSecret(ctx); err != nil {
|
||||
@@ -93,16 +96,17 @@ func applyReverseProxy(ctx context.Context, panelHost, adminHost string) error {
|
||||
// writeConnectionConfig stamps the chosen hostnames (and optional Access audience)
|
||||
// into both the host and pod config files. An empty aud clears any prior
|
||||
// Cloudflare audience, which is correct when switching to a non-Access front.
|
||||
func writeConnectionConfig(panelHost, adminHost, aud string) error {
|
||||
// clientIPHeader is the header that front writes the visitor address into.
|
||||
func writeConnectionConfig(panelHost, adminHost, aud, clientIPHeader string) error {
|
||||
for _, path := range []string{hostSetupConfigPath, podSetupConfigPath} {
|
||||
if err := updateAuthConfig(path, panelHost, adminHost, aud); err != nil {
|
||||
if err := updateAuthConfig(path, panelHost, adminHost, aud, clientIPHeader); err != nil {
|
||||
return err
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func updateAuthConfig(path, panelHost, adminHost, aud string) error {
|
||||
func updateAuthConfig(path, panelHost, adminHost, aud, clientIPHeader string) error {
|
||||
cfg, err := config.Load(path)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -112,6 +116,7 @@ func updateAuthConfig(path, panelHost, adminHost, aud string) error {
|
||||
}
|
||||
cfg.Auth.AdminHostname = adminHost
|
||||
cfg.Auth.AccessJWTAud = aud
|
||||
cfg.Auth.ClientIPHeader = clientIPHeader
|
||||
return writeConfig(path, cfg)
|
||||
}
|
||||
|
||||
@@ -133,6 +138,11 @@ func writeConfig(path string, cfg *config.Config) error {
|
||||
return os.Rename(tmpPath, path)
|
||||
}
|
||||
|
||||
// applyFelisConfigSecret applies the rendered config to the control namespace and
|
||||
// then converges the workload-namespace mirror best-effort. The mirror feeds the
|
||||
// backup/restore/fileedit Jobs and the reaper; without this refresh a reconfigure
|
||||
// here would leave those readers on the previous render until the next `felis
|
||||
// setup` run (startup pass) or installer re-run.
|
||||
func applyFelisConfigSecret(ctx context.Context) error {
|
||||
out, err := kubectlOutput(ctx,
|
||||
"-n", "felis", "create", "secret", "generic", "felis-config",
|
||||
@@ -142,7 +152,41 @@ func applyFelisConfigSecret(ctx context.Context) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return kubectlWithInput(ctx, out, "apply", "-f", "-")
|
||||
if err := kubectlWithInput(ctx, out, "apply", "-f", "-"); err != nil {
|
||||
return err
|
||||
}
|
||||
if err := replicateFelisConfigToWorkloadNamespace(ctx); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "felis setup: warning: the control-plane config is applied, but the workload-namespace mirror could not be refreshed (%v); re-run felis setup once that is fixed\n", err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// replicateFelisConfigToWorkloadNamespace overwrites the workload-namespace
|
||||
// felis-config mirror with the freshly rendered pod config. Deliberately a full
|
||||
// replace, not create-if-absent: a stale mirror is exactly what silently hands
|
||||
// the Jobs that mount it old settings after a reconfigure. No-op when the
|
||||
// workload namespace is unset or is the control namespace itself.
|
||||
func replicateFelisConfigToWorkloadNamespace(ctx context.Context) error {
|
||||
cfg, err := config.Load(hostSetupConfigPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ns := cfg.K8s.Namespace
|
||||
if ns == "" || ns == "felis" {
|
||||
return nil
|
||||
}
|
||||
manifest, err := kubectlOutput(ctx,
|
||||
"-n", ns, "create", "secret", "generic", "felis-config",
|
||||
"--from-file=felis.toml="+podSetupConfigPath,
|
||||
"--dry-run=client", "-o", "yaml",
|
||||
)
|
||||
if err != nil {
|
||||
return fmt.Errorf("render felis-config for %s: %w", ns, err)
|
||||
}
|
||||
if err := kubectlWithInput(ctx, manifest, "-n", ns, "apply", "-f", "-"); err != nil {
|
||||
return fmt.Errorf("replicate felis-config to %s: %w", ns, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func installCloudflaredService(ctx context.Context, cloudflaredBin, configPath string) error {
|
||||
|
||||
@@ -207,7 +207,11 @@ func (m *mcBindModel) doneView() string {
|
||||
if box.Len() > 0 {
|
||||
box.WriteString("\n")
|
||||
}
|
||||
box.WriteString(tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n")
|
||||
box.WriteString(tuiLabel.Render("setup URL ") + "\n")
|
||||
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
|
||||
box.WriteString(tuiPassword.Render(line) + "\n")
|
||||
}
|
||||
box.WriteString("\n")
|
||||
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup.\nIt is shown only once."))
|
||||
}
|
||||
if m.auditWarning != "" {
|
||||
|
||||
@@ -151,19 +151,46 @@ func TestOwnerModelProvisionErrorRouting(t *testing.T) {
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("a conflict on the Owner path is not a retry", func(t *testing.T) {
|
||||
// Defensive: the Owner upserts and so never conflicts, but were one ever to
|
||||
// surface it must end the session rather than loop the form — only the
|
||||
// insert-only operator path is retryable.
|
||||
t.Run("the Owner seat refusal returns to the form naming the seat", func(t *testing.T) {
|
||||
// Upserting a fresh username while a seat is occupied would mint a second
|
||||
// owner, so provisionOwner refuses with ownerSeatTakenError (Is
|
||||
// api.ErrConflict) and the console must route back for a retype — the same
|
||||
// recoverable contract as the operator clash, and the only Owner-path
|
||||
// conflict there is.
|
||||
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
|
||||
seatErr := &ownerSeatTakenError{seat: "seat-holder"}
|
||||
|
||||
next, cmd := m.Update(owProvisionMsg{err: conflict})
|
||||
next, cmd := m.Update(owProvisionMsg{err: seatErr})
|
||||
om := next.(*ownerModel)
|
||||
if om.step != owProvision {
|
||||
t.Fatalf("step = %v, want owProvision — the seat refusal is recoverable", om.step)
|
||||
}
|
||||
if om.provisionErr == nil || !errors.Is(om.provisionErr, api.ErrConflict) || !strings.Contains(om.provisionErr.Error(), "seat-holder") {
|
||||
t.Errorf("provisionErr = %v, want the seat refusal naming the seat", om.provisionErr)
|
||||
}
|
||||
// Feed the rebuilt form's init message back through so its view renders;
|
||||
// then the note must carry the seat name (the operator's retype cue).
|
||||
if cmd != nil {
|
||||
if msg := cmd(); msg != nil {
|
||||
if n2, _ := om.Update(msg); n2 != nil {
|
||||
om = n2.(*ownerModel)
|
||||
}
|
||||
}
|
||||
}
|
||||
if view := om.form.View(); !strings.Contains(view, "seat-holder") {
|
||||
t.Errorf("the provision form must surface the seat refusal:\n%s", view)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("a generic Owner-path fault still tears the console down", func(t *testing.T) {
|
||||
m := newOwnerModel(ctx, &fakeOwnerStore{}, "root", true)
|
||||
next, cmd := m.Update(owProvisionMsg{err: errors.New("boom")})
|
||||
om := next.(*ownerModel)
|
||||
if om.provisionErr != nil {
|
||||
t.Error("the Owner path recorded a retryable conflict; only the operator path retries")
|
||||
t.Error("a generic fault must not be treated as a retryable refusal")
|
||||
}
|
||||
if res, ok := cmd().(ownerResultMsg); !ok || res.err == nil {
|
||||
t.Error("an Owner-path conflict should tear down via an error result")
|
||||
t.Error("a generic Owner-path fault should tear down via an error result")
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
@@ -3,8 +3,10 @@ package main
|
||||
import (
|
||||
"context"
|
||||
"fmt"
|
||||
"io"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/dbbackup"
|
||||
"felis.lolicon.best/internal/store"
|
||||
)
|
||||
|
||||
@@ -12,7 +14,7 @@ import (
|
||||
// applied count. Used by the preflight stage to self-heal a freshly bootstrapped
|
||||
// (or upgraded) database.
|
||||
func applyMigrations(dbURL string) (int, error) {
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Minute)
|
||||
defer cancel()
|
||||
drv, err := store.Open(ctx, dbURL)
|
||||
if err != nil {
|
||||
@@ -23,6 +25,11 @@ func applyMigrations(dbURL string) (int, error) {
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
// Same guard as `felis migrate up`: never roll a populated database forward
|
||||
// without a snapshot to roll back to.
|
||||
if _, err := preMigrateBackup(ctx, drv, migrations, dbURL, dbbackup.DefaultDir, io.Discard); err != nil {
|
||||
return 0, fmt.Errorf("pre-migration backup: %w", err)
|
||||
}
|
||||
if _, err := store.Up(ctx, drv, migrations); err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
+21
-10
@@ -174,12 +174,13 @@ func (m *ownerModel) Update(msg tea.Msg) (tea.Model, tea.Cmd) {
|
||||
|
||||
case owProvisionMsg:
|
||||
if msg.err != nil {
|
||||
// A taken Operator username is the expected, recoverable outcome of the
|
||||
// insert-only operator path (refusing the clash is the whole reason it is
|
||||
// insert-only, not an upsert). Route back to the form with a note so the
|
||||
// operator can pick another name, rather than tearing down the console —
|
||||
// any other error is a genuine fault and still ends the session.
|
||||
if m.operation == bgAddOperator && errors.Is(msg.err, api.ErrConflict) {
|
||||
// api.ErrConflict marks the two recoverable refusals: a taken Operator
|
||||
// username (insert-only clash) and an Owner reset naming anything but the
|
||||
// occupied seat (ownerSeatTakenError Is ErrConflict). Route back to the
|
||||
// form with a note so the operator can retype, rather than tearing down
|
||||
// the console — any other error is a genuine fault and still ends the
|
||||
// session.
|
||||
if errors.Is(msg.err, api.ErrConflict) {
|
||||
m.provisionErr = msg.err
|
||||
m.step = owProvision
|
||||
m.form = m.sized(m.buildProvisionForm())
|
||||
@@ -364,9 +365,15 @@ func (m *ownerModel) buildProvisionForm() *huh.Form {
|
||||
}
|
||||
}
|
||||
if m.provisionErr != nil {
|
||||
// The only error routed back to this form is a username clash on the insert-only
|
||||
// operator path; show a concrete prompt to choose another name.
|
||||
desc = "That username is already taken — choose a different one.\n\n" + desc
|
||||
// Recoverable refusals routed back here: the seat refusal already names the
|
||||
// username to enter, so show it verbatim; the operator-name clash gets the
|
||||
// generic retry prompt.
|
||||
note := "That username is already taken — choose a different one."
|
||||
var seatErr *ownerSeatTakenError
|
||||
if errors.As(m.provisionErr, &seatErr) {
|
||||
note = seatErr.Error()
|
||||
}
|
||||
desc = note + "\n\n" + desc
|
||||
}
|
||||
|
||||
fields := []huh.Field{
|
||||
@@ -420,7 +427,11 @@ func (m *ownerModel) doneView() string {
|
||||
var box strings.Builder
|
||||
box.WriteString(tuiLabel.Render("username ") + m.username + "\n")
|
||||
if m.setupTokenURL != "" {
|
||||
box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n" + tuiPassword.Render(m.setupTokenURL) + "\n\n")
|
||||
box.WriteString("\n" + tuiLabel.Render("setup URL ") + "\n")
|
||||
for _, line := range wrapDisplayURL(m.setupTokenURL, 70) {
|
||||
box.WriteString(tuiPassword.Render(line) + "\n")
|
||||
}
|
||||
box.WriteString("\n")
|
||||
box.WriteString(tuiWarn.Render("Open this URL to complete passwordless login setup. It is shown only once."))
|
||||
}
|
||||
if m.auditWarning != "" {
|
||||
|
||||
@@ -553,6 +553,7 @@ func (m *rootModel) showSummary() (tea.Model, tea.Cmd) {
|
||||
storageLabel: m.result.storageDetail,
|
||||
routedHosts: routed,
|
||||
localHint: m.result.connectMethod == connectLocal,
|
||||
alreadySetUp: m.result.alreadySetUp,
|
||||
})
|
||||
}
|
||||
|
||||
|
||||
@@ -249,6 +249,34 @@ func TestRootReconfigureSMTP(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestRootReconfigureStorageKeepsStatusFraming locks the same rule for the
|
||||
// "change storage" path: on a re-run, completing it must land back on the
|
||||
// alreadySetUp status framing (with the updated recap), not "Setup complete."
|
||||
func TestRootReconfigureStorageKeepsStatusFraming(t *testing.T) {
|
||||
m := newTestRoot(true, consoleModeSetup, "")
|
||||
m = drive(t, m, preflightDoneMsg{})
|
||||
if _, ok := m.screen.(*summaryModel); !ok {
|
||||
t.Fatalf("re-run after preflight, screen = %T, want *summaryModel", m.screen)
|
||||
}
|
||||
|
||||
m = drive(t, m, reconfigureStorageMsg{})
|
||||
if _, ok := m.screen.(*storageChooserModel); !ok {
|
||||
t.Fatalf("reconfigure-storage screen = %T, want *storageChooserModel", m.screen)
|
||||
}
|
||||
|
||||
m = drive(t, m, storageResultMsg{method: storageLocal, detail: "local disk · /var/lib/felis/uploads"})
|
||||
sum, ok := m.screen.(*summaryModel)
|
||||
if !ok {
|
||||
t.Fatalf("after reconfigure-storage, screen = %T, want *summaryModel", m.screen)
|
||||
}
|
||||
if !sum.alreadySetUp {
|
||||
t.Fatalf("after reconfigure-storage, summary should keep the alreadySetUp framing")
|
||||
}
|
||||
if sum.storageLabel != "local disk · /var/lib/felis/uploads" {
|
||||
t.Fatalf("storageLabel = %q, want the updated recap", sum.storageLabel)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRootRerunLandsOnStatus(t *testing.T) {
|
||||
// adminExists at start of a setup run = re-run: preflight should skip straight
|
||||
// to the "manage in panel" status screen, never touching owner/connect.
|
||||
|
||||
+51
-5
@@ -4,6 +4,7 @@ import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
@@ -338,20 +339,30 @@ func applySMTPConfig(ctx context.Context, in smtpInputs) error {
|
||||
if err := applyFelisConfigSecret(ctx); err != nil {
|
||||
return err
|
||||
}
|
||||
// Refresh the workload-namespace copies too (the reaper's warning path): the
|
||||
// OTP path is already live in the control namespace, so a replica miss is
|
||||
// reported but not fatal.
|
||||
if err := replicateSMTPToWorkloadNamespace(ctx, in.password); err != nil {
|
||||
fmt.Fprintf(os.Stderr, "felis setup: warning: email is configured, but refreshing the workload copies failed (pre-reap warning emails may stay suppressed): %v\n", err)
|
||||
}
|
||||
if err := kubectl(ctx, "-n", "felis", "rollout", "restart", "deployment/felis-api"); err != nil {
|
||||
return err
|
||||
}
|
||||
return kubectl(ctx, "-n", "felis", "rollout", "status", "deployment/felis-api", "--timeout=180s")
|
||||
}
|
||||
|
||||
// applySMTPSecret creates (or replaces) the felis-smtp Secret the felis-api
|
||||
// Deployment injects the relay password from. Rendered in-process and piped to
|
||||
// smtpSecretManifest renders the felis-smtp Secret for the given namespace, the
|
||||
// one the receiving Deployment/CronJob resolves its secretKeyRef against (felis
|
||||
// for felis-api, the workload namespace for the reaper's mirror). The namespace
|
||||
// must be IN the manifest: kubectl rejects a manifest whose namespace conflicts
|
||||
// with -n, so leaving the control namespace hardcoded made every workload-ns
|
||||
// replica fail before it started. Rendered in-process and piped to
|
||||
// `kubectl apply` — the password is never a command-line arg, so it never
|
||||
// appears in the host process table.
|
||||
func applySMTPSecret(ctx context.Context, password string) error {
|
||||
func smtpSecretManifest(password, namespace string) ([]byte, error) {
|
||||
secret := &corev1.Secret{
|
||||
TypeMeta: metav1.TypeMeta{APIVersion: "v1", Kind: "Secret"},
|
||||
ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: "felis"},
|
||||
ObjectMeta: metav1.ObjectMeta{Name: platform.SMTPSecretName, Namespace: namespace},
|
||||
Type: corev1.SecretTypeOpaque,
|
||||
StringData: map[string]string{
|
||||
platform.SMTPSecretPasswordKey: password,
|
||||
@@ -359,7 +370,42 @@ func applySMTPSecret(ctx context.Context, password string) error {
|
||||
}
|
||||
manifest, err := yaml.Marshal(secret)
|
||||
if err != nil {
|
||||
return fmt.Errorf("render smtp secret: %w", err)
|
||||
return nil, fmt.Errorf("render smtp secret: %w", err)
|
||||
}
|
||||
return manifest, nil
|
||||
}
|
||||
|
||||
func applySMTPSecret(ctx context.Context, password string) error {
|
||||
manifest, err := smtpSecretManifest(password, "felis")
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
return kubectlWithInput(ctx, manifest, "apply", "-f", "-")
|
||||
}
|
||||
|
||||
// replicateSMTPToWorkloadNamespace refreshes the workload-namespace (minecraft)
|
||||
// copy of felis-smtp after email is reconfigured. The reaper's CronJob runs
|
||||
// there and resolves the password by local reference — a secretKeyRef is
|
||||
// namespace-local — so without this refresh a later SMTP change would never
|
||||
// reach the pre-reap warning emails. Deliberately OVERWRITES: this is a mirror
|
||||
// of the control-namespace source, and a stale mirror is exactly the failure
|
||||
// this closes. The felis-config mirror rides along in applyFelisConfigSecret,
|
||||
// which every apply path refreshes.
|
||||
func replicateSMTPToWorkloadNamespace(ctx context.Context, password string) error {
|
||||
cfg, err := config.Load(hostSetupConfigPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
ns := cfg.K8s.Namespace
|
||||
if ns == "" || ns == "felis" {
|
||||
return nil
|
||||
}
|
||||
smtpManifest, err := smtpSecretManifest(password, ns)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
if err := kubectlWithInput(ctx, smtpManifest, "-n", ns, "apply", "-f", "-"); err != nil {
|
||||
return fmt.Errorf("replicate %s to %s: %w", platform.SMTPSecretName, ns, err)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
@@ -0,0 +1,37 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sigs.k8s.io/yaml"
|
||||
)
|
||||
|
||||
// TestSMTPSecretManifestCarriesTargetNamespace pins the fix for the
|
||||
// workload-namespace replica: kubectl refuses a manifest whose namespace
|
||||
// conflicts with -n ("the namespace from the provided object ... does not
|
||||
// match"), so the mirror must render felis-smtp with the TARGET namespace —
|
||||
// otherwise the "configure email" refresh fails on the first apply and the
|
||||
// felis-config mirror never runs at all.
|
||||
func TestSMTPSecretManifestCarriesTargetNamespace(t *testing.T) {
|
||||
for _, ns := range []string{"felis", "minecraft"} {
|
||||
b, err := smtpSecretManifest("pw", ns)
|
||||
if err != nil {
|
||||
t.Fatalf("render for %s: %v", ns, err)
|
||||
}
|
||||
var got struct {
|
||||
Metadata struct {
|
||||
Namespace string `json:"namespace"`
|
||||
} `json:"metadata"`
|
||||
}
|
||||
if err := yaml.Unmarshal(b, &got); err != nil {
|
||||
t.Fatalf("unmarshal for %s: %v", ns, err)
|
||||
}
|
||||
if got.Metadata.Namespace != ns {
|
||||
t.Fatalf("manifest namespace = %q, want %q", got.Metadata.Namespace, ns)
|
||||
}
|
||||
if !strings.Contains(string(b), "name: felis-smtp") {
|
||||
t.Fatalf("manifest must still name felis-smtp: %s", b)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -29,6 +29,33 @@ func tuiSeparator() string {
|
||||
return tuiHint.Render(strings.Repeat("─", 70))
|
||||
}
|
||||
|
||||
// wrapDisplayURL breaks a long URL into lines no wider than width so the TUI
|
||||
// renderer never truncates it on a narrow terminal — the one-time setup URL
|
||||
// carries a 43-char token and overruns 80 columns. It prefers breaking right
|
||||
// after a '=' or '/' inside the window (the token then lands on its own line)
|
||||
// and hard-wraps only when no boundary is available. Lines concatenate back to
|
||||
// the original string.
|
||||
func wrapDisplayURL(u string, width int) []string {
|
||||
if width <= 0 {
|
||||
width = 70
|
||||
}
|
||||
var lines []string
|
||||
for len(u) > width {
|
||||
cut := width
|
||||
if i := strings.LastIndexByte(u[:width], '='); i >= 0 && i >= width/2 {
|
||||
cut = i + 1
|
||||
} else if i := strings.LastIndexByte(u[:width], '/'); i >= 0 && i >= width/2 {
|
||||
cut = i + 1
|
||||
}
|
||||
lines = append(lines, u[:cut])
|
||||
u = u[cut:]
|
||||
}
|
||||
if u != "" {
|
||||
lines = append(lines, u)
|
||||
}
|
||||
return lines
|
||||
}
|
||||
|
||||
// tuiStepRail renders a breadcrumb of wizard stages. Steps before `current`
|
||||
// render as done, `current` is highlighted, and later steps are dimmed.
|
||||
func tuiStepRail(steps []string, current int) string {
|
||||
|
||||
@@ -0,0 +1,26 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestWrapDisplayURL(t *testing.T) {
|
||||
u := "https://op.console.example.net/setup?token=" + strings.Repeat("A", 43)
|
||||
lines := wrapDisplayURL(u, 70)
|
||||
if got := strings.Join(lines, ""); got != u {
|
||||
t.Fatalf("concatenated lines = %q, want the original URL back", got)
|
||||
}
|
||||
for i, l := range lines {
|
||||
if len(l) > 70 {
|
||||
t.Errorf("line %d is %d cols wide: %q", i, len(l), l)
|
||||
}
|
||||
}
|
||||
if len(lines) < 2 || !strings.HasSuffix(lines[0], "token=") {
|
||||
t.Fatalf("want the first line to end at the 'token=' boundary, got %q", lines)
|
||||
}
|
||||
short := "https://a/b"
|
||||
if got := wrapDisplayURL(short, 70); len(got) != 1 || got[0] != short {
|
||||
t.Errorf("short URL should pass through unsplit, got %q", got)
|
||||
}
|
||||
}
|
||||
+62
-27
@@ -24,8 +24,9 @@ const updateTimeout = 60 * time.Second
|
||||
// The apply side is deliberately NOT implemented in this command. Every component
|
||||
// here is installed by deploy/bootstrap.sh, which is idempotent, already handles the
|
||||
// parts that are easy to get wrong (Velocity's pinned MINOR, the atomic jar install,
|
||||
// the k3s image re-import that a byte-identical StatefulSet template will not
|
||||
// trigger on its own), and is the path that gets exercised on every install. A
|
||||
// the image re-import + registry push that a byte-identical StatefulSet template
|
||||
// will not trigger on its own), and is the path that gets exercised on every
|
||||
// install. A
|
||||
// second installer living in this file would duplicate that policy, could drift from
|
||||
// it silently, and would be reachable only on a live node where a mistake takes the
|
||||
// proxy or the control plane down. So `felis update` reports, and hands the operator
|
||||
@@ -42,6 +43,21 @@ type updateTarget struct {
|
||||
command string
|
||||
}
|
||||
|
||||
// installerRerun is the tested apply path for every planner-backed selector: re-run the
|
||||
// installer. It is idempotent, and it is the only path that fetches a newer version --
|
||||
// `felis setup` skips its host-bootstrap phase on a completed install (all four install
|
||||
// markers already exist), so there it opens the config console and moves no component,
|
||||
// and even on the bootstrap path it re-images felis-api from the binary setup is already
|
||||
// running (FELIS_BOOTSTRAP_BINARY), which looks like an update and changes nothing.
|
||||
//
|
||||
// The URL is the one-liner both READMEs hand out, read at a tag rather than main: the
|
||||
// script's release channel installs the newest release's binary, and main can carry
|
||||
// installer changes that binary was never tested with. installerRef picks the tag and
|
||||
// renderApplyGuidance substitutes it for {ref}. While the repo is private the URL
|
||||
// answers 404 (raw.githubusercontent.com hides private repos), which is why the trailer
|
||||
// below points at the README's token'd form for that case.
|
||||
const installerRerun = "curl -fsSL https://raw.githubusercontent.com/FelisMC/Felis/{ref}/deploy/bootstrap.sh | sudo bash"
|
||||
|
||||
// updateTargets is the selector table. panel and plugins both resolve to felis-api
|
||||
// because they are not separately versioned: the panel is compiled into the felis
|
||||
// binary with //go:embed, and the plugin jars are built from this same repo in the
|
||||
@@ -51,19 +67,19 @@ var updateTargets = []updateTarget{
|
||||
selector: "panel",
|
||||
component: "felis-api",
|
||||
note: "the panel is embedded in the felis binary (//go:embed), so updating it means rebuilding the felis image and rolling felis-api",
|
||||
command: "sudo felis setup",
|
||||
command: installerRerun,
|
||||
},
|
||||
{
|
||||
selector: "velocity",
|
||||
component: "velocity",
|
||||
note: "re-runs install_velocity: newest BUILD of the pinned minor (FELIS_VELOCITY_VERSION), atomic jar install, then restarts felis-velocity",
|
||||
command: "sudo felis setup",
|
||||
note: "re-runs install_velocity: the build the release pins in deploy/game-stack.lock (FELIS_VELOCITY_VERSION=<minor> takes that minor's newest build instead), sha256-checked, atomic jar install, then restarts felis-velocity only if the jar or its config changed",
|
||||
command: installerRerun,
|
||||
},
|
||||
{
|
||||
selector: "plugins",
|
||||
component: "felis-api",
|
||||
note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + k3s image re-import",
|
||||
command: "sudo felis setup",
|
||||
note: "felis-velocity.jar is a host-file swap, but felis-paper.jar and felis-limbo.jar are baked into the lobby/limbo images and need a rebuild + re-mirror into the in-cluster registry (the installer re-run does both)",
|
||||
command: installerRerun,
|
||||
},
|
||||
{
|
||||
selector: "mc",
|
||||
@@ -219,7 +235,6 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
|
||||
|
||||
var b strings.Builder
|
||||
var offeredCommand bool
|
||||
var offeredFelisAPI bool
|
||||
for _, t := range updateTargets {
|
||||
if !selected[t.selector] {
|
||||
continue
|
||||
@@ -246,30 +261,50 @@ func renderApplyGuidance(res updater.Result, selected map[string]bool, force boo
|
||||
// component, and reinstalling the current release is a valid repair action.
|
||||
fmt.Fprintf(&b, " note: cannot tell whether %s is current — its latest version could not be discovered (see above); this reinstalls it either way\n", t.component)
|
||||
}
|
||||
fmt.Fprintf(&b, " run: %s\n", t.command)
|
||||
fmt.Fprintf(&b, " run: %s\n", strings.ReplaceAll(t.command, "{ref}", installerRef(byComponent)))
|
||||
offeredCommand = true
|
||||
offeredFelisAPI = offeredFelisAPI || t.component == "felis-api"
|
||||
}
|
||||
// Only explain the command when one was actually offered; a --mc-only run has
|
||||
// nothing to run and the trailer would be a non-sequitur.
|
||||
if offeredCommand {
|
||||
b.WriteString("\nfelis setup is idempotent and re-runs the installer that owns these components;\nit does not reinstall what is already current. Restart game servers afterwards.\n")
|
||||
}
|
||||
// Scoped to felis-api because it is the only component setup cannot move forward.
|
||||
// velocity is fine: install_velocity re-resolves the newest build of the pinned minor
|
||||
// on every run. But setup hands deploy/bootstrap.sh the binary it is itself running
|
||||
// (FELIS_BOOTSTRAP_BINARY), and that arm skips the release lookup entirely, so it
|
||||
// rebuilds the image and rolls the deployment from the SAME binary -- a run that looks
|
||||
// like a successful update and leaves the version unchanged.
|
||||
//
|
||||
// The installer is the only thing that moves felis-api. It is safe to point at now
|
||||
// that detect_node_ip reuses the installed root domain, so what is left to warn about
|
||||
// is the channel: FELIS_VERSION_BOOTSTRAP is not persisted anywhere and defaults to
|
||||
// release, so a bare re-run on a host tracking main quietly moves it onto releases.
|
||||
// That is a channel change, not a broken install, which is why it is one clause and
|
||||
// not a paragraph.
|
||||
if offeredFelisAPI {
|
||||
b.WriteString("\nfelis-api (panel, plugins) is the exception: setup re-images it from the felis binary\nalready on this host, so it cannot install a NEWER felis-api. Re-run the bootstrap\ninstaller for that -- it keeps this install's root domain. It does default to the\nrelease channel, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main.\n")
|
||||
// One trailer serves every selector now: setup is not an apply path at all on a
|
||||
// completed install (shouldRunHostBootstrapBeforeConfig only enters the host
|
||||
// bootstrap while an install marker is missing), so the installer re-run is the one
|
||||
// worked path for all three components and there is no per-component exception left
|
||||
// to scope. Two caveats stay because following the advice without them bites real
|
||||
// hosts: the channel is not persisted anywhere (a bare re-run on a main host quietly
|
||||
// moves it onto releases), and the private repo's one-liner needs the read token
|
||||
// back in the environment before it can resolve anything.
|
||||
if offeredCommand {
|
||||
b.WriteString("\nRe-running the installer applies everything above: it fetches the newest version on\nthe channel in effect and re-applies the bundle (release is the default). The channel\nis not persisted, so pass FELIS_VERSION_BOOTSTRAP=dev if this host tracks main. While\nthis repo is private, the one-liner above 404s without a token; the README's install\nsection has the token'd form that works. felis setup is not this path: on a completed\ninstall it opens the config console and installs nothing newer. Restart game servers\nafterwards.\n")
|
||||
}
|
||||
return b.String()
|
||||
}
|
||||
|
||||
// installerRef is the git ref the installer re-run reads bootstrap.sh from: the newest
|
||||
// stable felis release when the feed answered, which is the release that script then
|
||||
// installs; else the release this host runs; main only when neither is a release tag.
|
||||
func installerRef(byComponent map[string]updates.Action) string {
|
||||
a, ok := byComponent["felis-api"]
|
||||
if !ok {
|
||||
return "main"
|
||||
}
|
||||
if a.LatestKnown && isReleaseTag(a.Latest) {
|
||||
return a.Latest.String()
|
||||
}
|
||||
if isReleaseTag(a.Current) {
|
||||
return a.Current.String()
|
||||
}
|
||||
return "main"
|
||||
}
|
||||
|
||||
// isReleaseTag reports whether v was read from a stable vX.Y.Z tag, the only refs
|
||||
// release.yml publishes a binary for.
|
||||
func isReleaseTag(v updates.Version) bool {
|
||||
s := v.String()
|
||||
if !strings.HasPrefix(s, "v") || v.IsPrerelease() {
|
||||
return false
|
||||
}
|
||||
_, err := updates.Parse(s)
|
||||
return err == nil
|
||||
}
|
||||
+52
-25
@@ -107,7 +107,7 @@ func TestApplyGuidanceMinecraftOffersNoCommand(t *testing.T) {
|
||||
if !strings.Contains(out, "pinned by policy") {
|
||||
t.Fatalf("want the pin explained:\n%s", out)
|
||||
}
|
||||
if strings.Contains(out, "run:") || strings.Contains(out, "felis setup is idempotent") {
|
||||
if strings.Contains(out, "run:") || strings.Contains(out, "Re-running the installer") {
|
||||
t.Fatalf("--mc must offer no command and no command trailer:\n%s", out)
|
||||
}
|
||||
}
|
||||
@@ -133,42 +133,69 @@ func TestUpdateTargetsMatchTopology(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// `sudo felis setup` is the right answer for velocity and the wrong one for felis-api,
|
||||
// so the caveat has to be scoped rather than appended to every run. setup hands
|
||||
// bootstrap the binary it is already running, and that arm skips the release lookup:
|
||||
// the run rebuilds the image and rolls the deployment off the SAME binary, which looks
|
||||
// like a successful update and changes nothing. install_velocity, by contrast, really
|
||||
// does re-resolve the newest build on every run.
|
||||
func TestApplyGuidanceScopesTheFelisAPICaveat(t *testing.T) {
|
||||
const caveat = "cannot install a NEWER felis-api"
|
||||
|
||||
// Re-running the installer is the one apply path this table may hand out. setup is NOT an
|
||||
// updater on a completed install -- its host-bootstrap phase only runs while an install
|
||||
// marker is missing, so it opens the config console and moves no component -- and even on
|
||||
// the bootstrap path it re-images felis-api from the binary setup is already running. The
|
||||
// table used to answer with "sudo felis setup" and scope a felis-api-only exception; both
|
||||
// taught a model that does not survive contact with an installed host.
|
||||
func TestApplyGuidancePointsEveryComponentAtTheInstaller(t *testing.T) {
|
||||
api := renderApplyGuidance(
|
||||
planResult([]updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, LatestKnown: true}}),
|
||||
map[string]bool{"panel": true}, false)
|
||||
if !strings.Contains(api, caveat) {
|
||||
t.Fatalf("--panel resolves to felis-api and must carry the caveat:\n%s", api)
|
||||
for _, want := range []string{"deploy/bootstrap.sh", "FELIS_VERSION_BOOTSTRAP=dev", "felis setup is not this path"} {
|
||||
if !strings.Contains(api, want) {
|
||||
t.Fatalf("--panel guidance missing %q:\n%s", want, api)
|
||||
}
|
||||
}
|
||||
// Naming the installer obliges us to name what a bare re-run still changes. The domain
|
||||
// is handled -- detect_node_ip reuses the installed one -- but the channel is not
|
||||
// persisted at all and defaults to release, so a host tracking main gets moved onto
|
||||
// releases by following this advice.
|
||||
if !strings.Contains(api, "FELIS_VERSION_BOOTSTRAP=dev") {
|
||||
t.Fatalf("pointing at the installer without the channel caveat misleads a dev host:\n%s", api)
|
||||
if strings.Contains(api, "run: sudo felis setup") {
|
||||
t.Fatalf("setup must never be offered as the apply command:\n%s", api)
|
||||
}
|
||||
|
||||
// The same path serves velocity; a scoped caveat would re-teach the old model that
|
||||
// setup fixes velocity.
|
||||
vel := renderApplyGuidance(
|
||||
planResult([]updates.Action{{Component: "velocity", Kind: updates.ActionNotify, LatestKnown: true}}),
|
||||
map[string]bool{"velocity": true}, false)
|
||||
if strings.Contains(vel, caveat) {
|
||||
t.Fatalf("velocity IS fixed by setup; the caveat would misdirect the operator:\n%s", vel)
|
||||
}
|
||||
if !strings.Contains(vel, "felis setup is idempotent") {
|
||||
t.Fatalf("velocity still wants the ordinary trailer:\n%s", vel)
|
||||
if !strings.Contains(vel, "run: curl -fsSL") || !strings.Contains(vel, "felis setup is not this path") {
|
||||
t.Fatalf("velocity gets the same installer path:\n%s", vel)
|
||||
}
|
||||
|
||||
// --mc offers no command at all, so neither trailer belongs.
|
||||
mc := renderApplyGuidance(planResult(nil), map[string]bool{"mc": true}, true)
|
||||
if strings.Contains(mc, caveat) {
|
||||
t.Fatalf("--mc offers no command; the caveat is a non-sequitur:\n%s", mc)
|
||||
if strings.Contains(mc, "deploy/bootstrap.sh") || strings.Contains(mc, "FELIS_VERSION_BOOTSTRAP") {
|
||||
t.Fatalf("--mc offers no command; the trailer is a non-sequitur:\n%s", mc)
|
||||
}
|
||||
}
|
||||
|
||||
// The re-run reads bootstrap.sh at the tag whose binary it installs. main can carry
|
||||
// installer changes no release was tested with.
|
||||
func TestApplyGuidanceReadsTheInstallerAtTheReleaseTag(t *testing.T) {
|
||||
v := func(s string) updates.Version {
|
||||
t.Helper()
|
||||
out, err := updates.Parse(s)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
return out
|
||||
}
|
||||
cases := []struct {
|
||||
name string
|
||||
api []updates.Action
|
||||
want string
|
||||
}{
|
||||
{"latest known", []updates.Action{{Component: "felis-api", Kind: updates.ActionNotify, Current: v("v1.3.0"), Latest: v("v1.4.0"), LatestKnown: true}}, "/FelisMC/Felis/v1.4.0/deploy/bootstrap.sh"},
|
||||
{"latest unknown", []updates.Action{{Component: "felis-api", Kind: updates.ActionNone, Current: v("v1.3.0")}}, "/FelisMC/Felis/v1.3.0/deploy/bootstrap.sh"},
|
||||
{"prerelease latest", []updates.Action{{Component: "felis-api", Kind: updates.ActionNone, Current: v("v1.3.0"), Latest: v("v1.4.0-rc.1"), LatestKnown: true}}, "/FelisMC/Felis/v1.3.0/deploy/bootstrap.sh"},
|
||||
{"nothing known", nil, "/FelisMC/Felis/main/deploy/bootstrap.sh"},
|
||||
}
|
||||
for _, c := range cases {
|
||||
out := renderApplyGuidance(planResult(c.api), map[string]bool{"velocity": true, "panel": true}, true)
|
||||
if !strings.Contains(out, c.want) {
|
||||
t.Errorf("%s: want %q in:\n%s", c.name, c.want, out)
|
||||
}
|
||||
if strings.Contains(out, "{ref}") {
|
||||
t.Errorf("%s: placeholder left in:\n%s", c.name, out)
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,269 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"flag"
|
||||
"fmt"
|
||||
"io"
|
||||
"net"
|
||||
"os"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"felis.lolicon.best/internal/config"
|
||||
"felis.lolicon.best/internal/mail"
|
||||
"felis.lolicon.best/internal/offsite"
|
||||
"felis.lolicon.best/internal/platform"
|
||||
"felis.lolicon.best/internal/store"
|
||||
"felis.lolicon.best/internal/watchdog"
|
||||
corev1 "k8s.io/api/core/v1"
|
||||
apierrors "k8s.io/apimachinery/pkg/api/errors"
|
||||
"sigs.k8s.io/controller-runtime/pkg/client"
|
||||
)
|
||||
|
||||
// proxyFor is how long the game proxy may refuse connections before it is
|
||||
// mailed: a restart takes seconds.
|
||||
const proxyFor = 3 * time.Minute
|
||||
|
||||
// cmdWatchdog runs one pass of the platform watchdog (internal/watchdog): it
|
||||
// checks the cluster, PostgreSQL, the game proxy, the database backups and the
|
||||
// host, prints every finding, and mails the platform owners what came due.
|
||||
// deploy/bootstrap.sh runs it every two minutes from felis-watchdog.timer.
|
||||
func cmdWatchdog(args []string, stdout, stderr io.Writer) int {
|
||||
fs := flag.NewFlagSet("watchdog", flag.ContinueOnError)
|
||||
fs.SetOutput(stderr)
|
||||
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml (the host copy, which reaches PostgreSQL on 127.0.0.1)")
|
||||
statePath := fs.String("state", "/var/lib/felis/watchdog/state.json", "state kept between runs (root only: it caches the relay password)")
|
||||
quietPath := fs.String("quiet-file", "/run/felis/watchdog-quiet-until", "Unix time before which nothing is mailed; the installer writes it while it restarts things on purpose")
|
||||
backupDir := fs.String("backup-dir", "/var/lib/felis/db-backups", `control-plane database backups to check for freshness ("" skips the check)`)
|
||||
diskPaths := fs.String("disk-paths", "/,/var/lib/rancher/k3s,/var/lib/postgresql,/var/lib/felis", "comma-separated paths whose filesystems must keep free space")
|
||||
proxyAddr := fs.String("proxy-addr", "", `game proxy address to dial, e.g. 127.0.0.1:25565 ("" skips the check)`)
|
||||
controlNS := fs.String("control-namespace", platform.DefaultControlNamespace, "namespace of the control plane")
|
||||
offsiteStatus := fs.String("offsite-status", offsite.DefaultStatusFile, "the record `felis offsite sync` leaves, checked when [offsite] is configured")
|
||||
toolsStatus := fs.String("build-tools-status", defaultBuildToolsStatus, "the record `felis mirror-build-tools` leaves, checked when builds scan against the registry's DB copy")
|
||||
dryRun := fs.Bool("dry-run", false, "print every finding and the mail that is due; send nothing and keep the state as it was")
|
||||
if err := fs.Parse(args); err != nil {
|
||||
if errors.Is(err, flag.ErrHelp) {
|
||||
return 0
|
||||
}
|
||||
return 2
|
||||
}
|
||||
cfg, err := config.Load(*cfgPath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis watchdog: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
state, err := watchdog.LoadState(*statePath)
|
||||
if err != nil {
|
||||
fmt.Fprintf(stderr, "felis watchdog: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
||||
defer cancel()
|
||||
now := time.Now()
|
||||
|
||||
var report watchdog.Report
|
||||
add := func(f *watchdog.Finding) {
|
||||
if f != nil {
|
||||
report.Findings = append(report.Findings, *f)
|
||||
}
|
||||
}
|
||||
|
||||
// The cluster: one unreachable API server stands in for every check behind it.
|
||||
minecraftNS := cfg.K8s.Namespace
|
||||
if minecraftNS == "" {
|
||||
minecraftNS = platform.DefaultMinecraftNamespace
|
||||
}
|
||||
cl, err := buildSystemServerClient()
|
||||
var found []watchdog.Finding
|
||||
if err == nil {
|
||||
found, err = watchdog.Cluster{Client: cl, ControlNamespace: *controlNS, MinecraftNamespace: minecraftNS}.Check(ctx, now)
|
||||
}
|
||||
if err != nil {
|
||||
f := watchdog.KubeAPIDown(err)
|
||||
add(&f)
|
||||
report.Unknown = append(report.Unknown, watchdog.ClusterPrefixes...)
|
||||
} else {
|
||||
report.Findings = append(report.Findings, found...)
|
||||
if cfg.SMTP.Host != "" {
|
||||
refreshSMTPPassword(ctx, cl, *controlNS, state, stderr)
|
||||
}
|
||||
}
|
||||
|
||||
if recipients, err := ownerEmails(ctx, cfg.Database.URL); err != nil {
|
||||
f := watchdog.PostgresDown(err)
|
||||
add(&f)
|
||||
} else {
|
||||
state.Recipients = recipients
|
||||
}
|
||||
|
||||
if *proxyAddr != "" {
|
||||
add(proxyFinding(ctx, *proxyAddr))
|
||||
}
|
||||
if *backupDir != "" {
|
||||
add(watchdog.BackupFinding(*backupDir, now))
|
||||
}
|
||||
if cfg.Offsite.Enabled() {
|
||||
add(watchdog.OffsiteFinding(*offsiteStatus, now))
|
||||
}
|
||||
if usesMirroredScanDB(cfg) {
|
||||
add(watchdog.ScanDBFinding(*toolsStatus, now))
|
||||
}
|
||||
report.Findings = append(report.Findings, watchdog.DiskFindings(splitList(*diskPaths))...)
|
||||
add(watchdog.MemoryFinding("/proc/meminfo"))
|
||||
|
||||
if len(report.Findings) == 0 {
|
||||
fmt.Fprintln(stdout, "felis watchdog: every check passed")
|
||||
}
|
||||
for _, f := range report.Findings {
|
||||
fmt.Fprintf(stdout, "felis watchdog: [%s] %s: %s\n", f.Severity, f.Key, f.SummaryEN)
|
||||
}
|
||||
|
||||
plan := state.Observe(report, now)
|
||||
host, _ := os.Hostname()
|
||||
subject, body := plan.Message(host, now)
|
||||
if *dryRun {
|
||||
if plan.Empty() {
|
||||
fmt.Fprintln(stdout, "felis watchdog: nothing is due to be mailed")
|
||||
} else {
|
||||
fmt.Fprintf(stdout, "felis watchdog: due to be mailed to %s:\nSubject: %s\n\n%s", strings.Join(state.Recipients, ", "), subject, strings.ReplaceAll(body, "\r\n", "\n"))
|
||||
}
|
||||
return 0
|
||||
}
|
||||
|
||||
save := func() int {
|
||||
if err := watchdog.SaveState(*statePath, state); err != nil {
|
||||
fmt.Fprintf(stderr, "felis watchdog: save state: %v\n", err)
|
||||
return 1
|
||||
}
|
||||
return 0
|
||||
}
|
||||
if plan.Empty() {
|
||||
return save()
|
||||
}
|
||||
if until := watchdog.QuietUntil(*quietPath); now.Before(until) {
|
||||
fmt.Fprintf(stdout, "felis watchdog: quiet until %s (installer running); holding this mail: %s\n", until.UTC().Format(time.RFC3339), subject)
|
||||
return save()
|
||||
}
|
||||
switch {
|
||||
case cfg.SMTP.Host == "":
|
||||
fmt.Fprintf(stdout, "felis watchdog: no [smtp] relay configured, so this is logged only: %s\n", subject)
|
||||
case len(state.Recipients) == 0:
|
||||
fmt.Fprintf(stdout, "felis watchdog: no owner account has a verified email, so this is logged only: %s\n", subject)
|
||||
default:
|
||||
if err := sendAlert(ctx, cfg, state, subject, body); err != nil {
|
||||
// Not committed: the same alerts come due again next run.
|
||||
fmt.Fprintf(stderr, "felis watchdog: mail %q: %v\n", subject, err)
|
||||
save()
|
||||
return 1
|
||||
}
|
||||
fmt.Fprintf(stdout, "felis watchdog: mailed %s: %s\n", strings.Join(state.Recipients, ", "), subject)
|
||||
}
|
||||
state.Commit(plan, now)
|
||||
return save()
|
||||
}
|
||||
|
||||
// usesMirroredScanDB reports whether build scans read the vulnerability DB copy
|
||||
// felis mirror-build-tools keeps in the platform registry: the default, or an
|
||||
// explicit trivy_db_repository under the registry's mirror/.
|
||||
func usesMirroredScanDB(cfg *config.Config) bool {
|
||||
if cfg.Registry.URL == "" {
|
||||
return false
|
||||
}
|
||||
repo := cfg.Registry.TrivyDBRepository
|
||||
return repo == "" || strings.HasPrefix(repo, cfg.Registry.URL+"/mirror/")
|
||||
}
|
||||
|
||||
// refreshSMTPPassword caches the relay password from the felis-smtp Secret, or
|
||||
// forgets it when the Secret is gone (a relay without AUTH). An env var named by
|
||||
// [smtp] password_ref, when set, wins at send time instead.
|
||||
func refreshSMTPPassword(ctx context.Context, cl client.Client, ns string, state *watchdog.State, stderr io.Writer) {
|
||||
var sec corev1.Secret
|
||||
err := cl.Get(ctx, client.ObjectKey{Namespace: ns, Name: platform.SMTPSecretName}, &sec)
|
||||
switch {
|
||||
case apierrors.IsNotFound(err):
|
||||
state.SMTPPassword = ""
|
||||
case err != nil:
|
||||
fmt.Fprintf(stderr, "felis watchdog: read %s/%s (keeping the cached relay password): %v\n", ns, platform.SMTPSecretName, err)
|
||||
default:
|
||||
state.SMTPPassword = string(sec.Data[platform.SMTPSecretPasswordKey])
|
||||
}
|
||||
}
|
||||
|
||||
// ownerEmails pings PostgreSQL and returns the verified addresses of the
|
||||
// enabled owner accounts, the people who can act on an alert.
|
||||
func ownerEmails(ctx context.Context, url string) ([]string, error) {
|
||||
ctx, cancel := context.WithTimeout(ctx, 15*time.Second)
|
||||
defer cancel()
|
||||
drv, err := store.Open(ctx, url)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer drv.Close()
|
||||
rows, err := drv.DB().QueryContext(ctx,
|
||||
`SELECT email FROM users
|
||||
WHERE role = 'owner' AND email_verified AND COALESCE(email, '') <> ''
|
||||
AND NOT disabled AND deleted_at IS NULL
|
||||
ORDER BY email`)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
defer rows.Close()
|
||||
var out []string
|
||||
for rows.Next() {
|
||||
var email sql.NullString
|
||||
if err := rows.Scan(&email); err != nil {
|
||||
return nil, err
|
||||
}
|
||||
out = append(out, email.String)
|
||||
}
|
||||
return out, rows.Err()
|
||||
}
|
||||
|
||||
// proxyFinding dials the game proxy; players reach every server through it.
|
||||
func proxyFinding(ctx context.Context, addr string) *watchdog.Finding {
|
||||
d := net.Dialer{Timeout: 5 * time.Second}
|
||||
conn, err := d.DialContext(ctx, "tcp", addr)
|
||||
if err == nil {
|
||||
conn.Close()
|
||||
return nil
|
||||
}
|
||||
return &watchdog.Finding{
|
||||
Key: "proxy", Severity: watchdog.Critical, For: proxyFor,
|
||||
Summary: fmt.Sprintf("游戏代理 %s 无法连接:玩家进不了任何服务器", addr),
|
||||
SummaryEN: fmt.Sprintf("the game proxy at %s refuses connections: players cannot reach any server", addr),
|
||||
Hint: fmt.Sprintf("systemctl status felis-velocity; journalctl -u felis-velocity -n 200 (%v)", err),
|
||||
}
|
||||
}
|
||||
|
||||
// sendAlert mails subject/body to every recipient; it fails only when no
|
||||
// recipient got it.
|
||||
func sendAlert(ctx context.Context, cfg *config.Config, state *watchdog.State, subject, body string) error {
|
||||
password := state.SMTPPassword
|
||||
if ref := cfg.SMTP.PasswordRef; ref != "" && os.Getenv(ref) != "" {
|
||||
password = os.Getenv(ref)
|
||||
}
|
||||
relay := &mail.SMTP{Host: cfg.SMTP.Host, Port: cfg.SMTP.Port, From: cfg.SMTP.From, Username: cfg.SMTP.Username, Password: password}
|
||||
var errs []error
|
||||
for _, to := range state.Recipients {
|
||||
if err := relay.SendNotice(ctx, to, subject, body); err != nil {
|
||||
errs = append(errs, fmt.Errorf("%s: %w", to, err))
|
||||
}
|
||||
}
|
||||
if len(errs) == len(state.Recipients) {
|
||||
return errors.Join(errs...)
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func splitList(s string) []string {
|
||||
var out []string
|
||||
for _, p := range strings.Split(s, ",") {
|
||||
if p = strings.TrimSpace(p); p != "" {
|
||||
out = append(out, p)
|
||||
}
|
||||
}
|
||||
return out
|
||||
}
|
||||
@@ -0,0 +1,42 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"context"
|
||||
"net"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// TestProxyFinding: a listening proxy is healthy; a closed port is the critical
|
||||
// "players cannot reach any server" finding.
|
||||
func TestProxyFinding(t *testing.T) {
|
||||
ln, err := net.Listen("tcp", "127.0.0.1:0")
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
addr := ln.Addr().String()
|
||||
go func() {
|
||||
for {
|
||||
c, err := ln.Accept()
|
||||
if err != nil {
|
||||
return
|
||||
}
|
||||
c.Close()
|
||||
}
|
||||
}()
|
||||
if f := proxyFinding(context.Background(), addr); f != nil {
|
||||
t.Fatalf("listening proxy reported: %+v", f)
|
||||
}
|
||||
ln.Close()
|
||||
f := proxyFinding(context.Background(), addr)
|
||||
if f == nil || f.Key != "proxy" || !strings.Contains(f.SummaryEN, addr) {
|
||||
t.Fatalf("closed proxy = %+v, want the proxy finding", f)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSplitList(t *testing.T) {
|
||||
got := splitList(" /, /var/lib/felis ,,")
|
||||
if strings.Join(got, "|") != "/|/var/lib/felis" {
|
||||
t.Fatalf("splitList = %q", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,284 @@
|
||||
# Felis alert rules — plain Prometheus format (also the promtool-tested source
|
||||
# for felis-prometheusrule.yaml). See docs/troubleshooting.md §14 for scraping
|
||||
# and loading instructions. These are for a deployment that brings its own
|
||||
# Prometheus; every install already runs `felis watchdog` on the host, which
|
||||
# checks the same conditions without one and mails the owners (§14).
|
||||
#
|
||||
# felis_* series come from these processes:
|
||||
# - felis-operator-metrics Service :8080 → felis_servers_total, felis_server_phase,
|
||||
# felis_start_duration_seconds,
|
||||
# felis_build_info{component="operator"},
|
||||
# controller_runtime_*, workqueue_*
|
||||
# - felis-api-internal Service :8081 → felis_image_build_failures_total,
|
||||
# felis_mail_total, felis_rate_limited_total,
|
||||
# felis_auth_otp_lockouts_total,
|
||||
# felis_auth_failures_total,
|
||||
# felis_audit_write_failures_total,
|
||||
# felis_build_info{component="api"}
|
||||
# - node-exporter textfile collector → felis_db_backup_* (felis-db-backup.timer)
|
||||
# node_* / kube_* series come from node-exporter / kube-state-metrics.
|
||||
groups:
|
||||
- name: felis.rules
|
||||
rules:
|
||||
- alert: FelisImageBuildFailures
|
||||
expr: increase(felis_image_build_failures_total[6h]) > 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "modpack/image build failed in the last 6h"
|
||||
description: >-
|
||||
felis_image_build_failures_total increased. Inspect the failed build Job
|
||||
(kubectl logs -n felis-build job/<build-job>); the same error text is on
|
||||
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
|
||||
- alert: FelisSlowServerStarts
|
||||
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "p90 server start time exceeds 5 minutes"
|
||||
description: >-
|
||||
Starts regularly take over five minutes (felis_start_duration_seconds,
|
||||
observed when readiness is first reached). A start that never completes
|
||||
records nothing — cross-check desiredState=Running servers with no ready
|
||||
phase (troubleshooting §1).
|
||||
- name: felis.platform.rules
|
||||
rules:
|
||||
- alert: FelisOperatorDown
|
||||
expr: absent(felis_build_info{component="operator"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisAPIDown
|
||||
expr: absent(felis_build_info{component="api"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisLoginGateDown
|
||||
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §1, §2).
|
||||
- alert: FelisSystemServerDown
|
||||
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- alert: FelisServerFailed
|
||||
expr: felis_server_phase{role="",phase="Failed"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "server {{ $labels.server }} is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §2).
|
||||
- alert: FelisReconcileErrors
|
||||
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- alert: FelisReconcileStuck
|
||||
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: felis.jobs.rules
|
||||
rules:
|
||||
- alert: FelisWorldJobFailed
|
||||
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.job_name }} failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
|
||||
(troubleshooting §10).
|
||||
- alert: FelisReaperStale
|
||||
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: felis.node.rules
|
||||
rules:
|
||||
- alert: FelisNodeDiskSpaceLow
|
||||
expr: >-
|
||||
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
|
||||
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
|
||||
description: >-
|
||||
Sustained disk pressure evicts game pods and garbage-collects images
|
||||
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
|
||||
- alert: FelisNodeDiskPressure
|
||||
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
|
||||
description: >-
|
||||
The eviction chain is in progress: control-plane pods hold
|
||||
system-cluster-critical and survive, game pods do not. Free disk now
|
||||
(troubleshooting §13b).
|
||||
- alert: FelisNodeMemoryLow
|
||||
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "node memory available below 10% for 15m"
|
||||
description: >-
|
||||
PostgreSQL, the control plane, the registry and game servers share one
|
||||
node; sustained memory pressure risks OOM kills.
|
||||
- name: felis.backup.rules
|
||||
rules:
|
||||
- alert: FelisDBBackupStale
|
||||
expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "no control-plane database backup in over 26h"
|
||||
description: >-
|
||||
felis-db-backup.timer runs daily; the newest bundle is more than a day
|
||||
old. Read `journalctl -u felis-db-backup` on the host, then take one now
|
||||
with `sudo felis db backup` (troubleshooting §16).
|
||||
- alert: FelisDBBackupMetricMissing
|
||||
expr: absent(felis_db_backup_last_success_timestamp_seconds)
|
||||
for: 2h
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "database backup freshness is not being scraped"
|
||||
description: >-
|
||||
No felis_db_backup_last_success_timestamp_seconds series, so
|
||||
FelisDBBackupStale cannot fire. Point node-exporter's
|
||||
--collector.textfile.directory at the directory of
|
||||
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
|
||||
(troubleshooting §16).
|
||||
- name: felis.auth.rules
|
||||
rules:
|
||||
- alert: FelisMailBudgetExhausted
|
||||
expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the install-wide mail budget refused mail"
|
||||
description: >-
|
||||
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
|
||||
spent, and every sign-in code is refused with 429 mail_rate_limited until
|
||||
it refills. Check felis_rate_limited_total for a flood before raising the
|
||||
budget (troubleshooting §17).
|
||||
- alert: FelisMailDeliveryFailing
|
||||
expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the SMTP relay refused mail in the last 15m"
|
||||
description: >-
|
||||
felis_mail_total{result="failed"} increased: sign-in codes are not being
|
||||
delivered (502 mail_undeliverable). The relay's reason is in the
|
||||
felis-api log (troubleshooting §17).
|
||||
- alert: FelisSignInFlood
|
||||
expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "sign-in doors refusing over 10 requests a minute"
|
||||
description: >-
|
||||
The per-address sign-in limit has been refusing callers for 10 minutes.
|
||||
A script is hammering the auth doors; if real users report rate_limited
|
||||
at once instead, [auth] client_ip_header is missing and everyone shares
|
||||
the proxy's address (troubleshooting §17).
|
||||
- alert: FelisOTPAccountLocked
|
||||
expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "an account's email-code sign-in locked after 10 wrong codes"
|
||||
description: >-
|
||||
Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}).
|
||||
The audit log names the account (action auth.otp.locked); the owner was
|
||||
mailed. Unless they fumbled codes, someone is guessing at it
|
||||
(troubleshooting §17).
|
||||
- alert: FelisSignInFailures
|
||||
expr: sum(increase(felis_auth_failures_total[15m])) > 30
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "over 30 refused sign-ins in 15 minutes"
|
||||
description: >-
|
||||
Wrong codes, unknown addresses or bad passkey assertions well above people
|
||||
mistyping: someone is guessing or enumerating. `sum by (door, reason)
|
||||
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
|
||||
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
|
||||
- alert: FelisAuditWriteFailing
|
||||
expr: increase(felis_audit_write_failures_total[15m]) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "felis-api failed to write audit rows"
|
||||
description: >-
|
||||
The actions went through but their audit rows were lost. The felis-api log
|
||||
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
|
||||
being unreachable or out of disk.
|
||||
@@ -0,0 +1,509 @@
|
||||
# promtool unit tests: `promtool test rules felis-alerts_test.yml`
|
||||
# Proves every shipped rule actually fires on its target condition (and stays
|
||||
# silent before it).
|
||||
rule_files:
|
||||
- felis-alerts.yaml
|
||||
evaluation_interval: 1m
|
||||
tests:
|
||||
- name: build failure and slow starts
|
||||
interval: 1m
|
||||
input_series:
|
||||
# counter: quiet for 5m, then one failure per step.
|
||||
- series: 'felis_image_build_failures_total'
|
||||
values: '0x5 1x15'
|
||||
# histogram: all observations land in the (300,600] bucket.
|
||||
- series: 'felis_start_duration_seconds_bucket{le="120"}'
|
||||
values: '0x22'
|
||||
- series: 'felis_start_duration_seconds_bucket{le="300"}'
|
||||
values: '0x22'
|
||||
- series: 'felis_start_duration_seconds_bucket{le="600"}'
|
||||
values: '0+10x21'
|
||||
- series: 'felis_start_duration_seconds_bucket{le="+Inf"}'
|
||||
values: '0+10x21'
|
||||
alert_rule_test:
|
||||
- eval_time: 2m
|
||||
alertname: FelisImageBuildFailures
|
||||
exp_alerts: []
|
||||
- eval_time: 20m
|
||||
alertname: FelisImageBuildFailures
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "modpack/image build failed in the last 6h"
|
||||
description: >-
|
||||
felis_image_build_failures_total increased. Inspect the failed build Job
|
||||
(kubectl logs -n felis-build job/<build-job>); the same error text is on
|
||||
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
|
||||
- eval_time: 20m
|
||||
alertname: FelisSlowServerStarts
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "p90 server start time exceeds 5 minutes"
|
||||
description: >-
|
||||
Starts regularly take over five minutes (felis_start_duration_seconds,
|
||||
observed when readiness is first reached). A start that never completes
|
||||
records nothing — cross-check desiredState=Running servers with no ready
|
||||
phase (troubleshooting §1).
|
||||
- name: node disk and memory thresholds
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'node_filesystem_avail_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
|
||||
values: '10x26'
|
||||
- series: 'node_filesystem_size_bytes{device="/dev/vda1",fstype="xfs",instance="node1",job="node-exporter",mountpoint="/"}'
|
||||
values: '100x26'
|
||||
- series: 'kube_node_status_condition{condition="DiskPressure",node="n1",status="true"}'
|
||||
values: '0x4 1x22'
|
||||
- series: 'node_memory_MemAvailable_bytes{instance="node1",job="node-exporter"}'
|
||||
values: '5x26'
|
||||
- series: 'node_memory_MemTotal_bytes{instance="node1",job="node-exporter"}'
|
||||
values: '100x26'
|
||||
alert_rule_test:
|
||||
- eval_time: 2m
|
||||
alertname: FelisNodeDiskPressure
|
||||
exp_alerts: []
|
||||
- eval_time: 20m
|
||||
alertname: FelisNodeDiskSpaceLow
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
device: /dev/vda1
|
||||
fstype: xfs
|
||||
instance: node1
|
||||
job: node-exporter
|
||||
mountpoint: /
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "node filesystem / below 15% available"
|
||||
description: >-
|
||||
Sustained disk pressure evicts game pods and garbage-collects images
|
||||
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
|
||||
- eval_time: 20m
|
||||
alertname: FelisNodeDiskPressure
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
condition: DiskPressure
|
||||
node: n1
|
||||
status: "true"
|
||||
severity: critical
|
||||
exp_annotations:
|
||||
summary: "kubelet reports DiskPressure on n1"
|
||||
description: >-
|
||||
The eviction chain is in progress: control-plane pods hold
|
||||
system-cluster-critical and survive, game pods do not. Free disk now
|
||||
(troubleshooting §13b).
|
||||
- eval_time: 20m
|
||||
alertname: FelisNodeMemoryLow
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
instance: node1
|
||||
job: node-exporter
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "node memory available below 10% for 15m"
|
||||
description: >-
|
||||
PostgreSQL, the control plane, the registry and game servers share one
|
||||
node; sustained memory pressure risks OOM kills.
|
||||
- name: database backup freshness
|
||||
interval: 1m
|
||||
input_series:
|
||||
# The newest bundle was taken at t=0 and none since.
|
||||
- series: 'felis_db_backup_last_success_timestamp_seconds{instance="node1",job="node-exporter",label="daily"}'
|
||||
values: '0x1630'
|
||||
alert_rule_test:
|
||||
- eval_time: 25h
|
||||
alertname: FelisDBBackupStale
|
||||
exp_alerts: []
|
||||
- eval_time: 27h
|
||||
alertname: FelisDBBackupStale
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
exp_annotations:
|
||||
summary: "no control-plane database backup in over 26h"
|
||||
description: >-
|
||||
felis-db-backup.timer runs daily; the newest bundle is more than a day
|
||||
old. Read `journalctl -u felis-db-backup` on the host, then take one now
|
||||
with `sudo felis db backup` (troubleshooting §16).
|
||||
- eval_time: 27h
|
||||
alertname: FelisDBBackupMetricMissing
|
||||
exp_alerts: []
|
||||
- name: database backup freshness not scraped
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'up{job="node-exporter"}'
|
||||
values: '1x200'
|
||||
alert_rule_test:
|
||||
- eval_time: 1h
|
||||
alertname: FelisDBBackupMetricMissing
|
||||
exp_alerts: []
|
||||
- eval_time: 3h
|
||||
alertname: FelisDBBackupMetricMissing
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "database backup freshness is not being scraped"
|
||||
description: >-
|
||||
No felis_db_backup_last_success_timestamp_seconds series, so
|
||||
FelisDBBackupStale cannot fire. Point node-exporter's
|
||||
--collector.textfile.directory at the directory of
|
||||
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
|
||||
(troubleshooting §16).
|
||||
- name: sign-in mail budget and relay
|
||||
interval: 1m
|
||||
input_series:
|
||||
# Created at zero on start; the budget refuses one mail at t=3m.
|
||||
- series: 'felis_mail_total{kind="otp",result="throttled",job="felis-api"}'
|
||||
values: '0 0 0 1x30'
|
||||
- series: 'felis_mail_total{kind="otp",result="failed",job="felis-api"}'
|
||||
values: '0x33'
|
||||
alert_rule_test:
|
||||
- eval_time: 2m
|
||||
alertname: FelisMailBudgetExhausted
|
||||
exp_alerts: []
|
||||
- eval_time: 5m
|
||||
alertname: FelisMailBudgetExhausted
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "the install-wide mail budget refused mail"
|
||||
description: >-
|
||||
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
|
||||
spent, and every sign-in code is refused with 429 mail_rate_limited until
|
||||
it refills. Check felis_rate_limited_total for a flood before raising the
|
||||
budget (troubleshooting §17).
|
||||
- eval_time: 5m
|
||||
alertname: FelisMailDeliveryFailing
|
||||
exp_alerts: []
|
||||
- name: sign-in flood
|
||||
interval: 1m
|
||||
input_series:
|
||||
# 30 refusals a minute from t=0; a lone refused script at 2/min stays quiet.
|
||||
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
|
||||
values: '0+30x40'
|
||||
alert_rule_test:
|
||||
- eval_time: 10m
|
||||
alertname: FelisSignInFlood
|
||||
exp_alerts: []
|
||||
- eval_time: 20m
|
||||
alertname: FelisSignInFlood
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "sign-in doors refusing over 10 requests a minute"
|
||||
description: >-
|
||||
The per-address sign-in limit has been refusing callers for 10 minutes.
|
||||
A script is hammering the auth doors; if real users report rate_limited
|
||||
at once instead, [auth] client_ip_header is missing and everyone shares
|
||||
the proxy's address (troubleshooting §17).
|
||||
- name: sign-in trickle stays quiet
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_rate_limited_total{scope="auth_door",job="felis-api"}'
|
||||
values: '0+2x40'
|
||||
alert_rule_test:
|
||||
- eval_time: 30m
|
||||
alertname: FelisSignInFlood
|
||||
exp_alerts: []
|
||||
- name: account email-code lock
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_auth_otp_lockouts_total{purpose="login_email",job="felis-api"}'
|
||||
values: '0 0 1x90'
|
||||
- series: 'felis_auth_otp_lockouts_total{purpose="op_login",job="felis-api"}'
|
||||
values: '0x92'
|
||||
alert_rule_test:
|
||||
- eval_time: 1m
|
||||
alertname: FelisOTPAccountLocked
|
||||
exp_alerts: []
|
||||
- eval_time: 10m
|
||||
alertname: FelisOTPAccountLocked
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
purpose: login_email
|
||||
exp_annotations:
|
||||
summary: "an account's email-code sign-in locked after 10 wrong codes"
|
||||
description: >-
|
||||
Someone entered 10 wrong codes for one account within 24h (login_email).
|
||||
The audit log names the account (action auth.otp.locked); the owner was
|
||||
mailed. Unless they fumbled codes, someone is guessing at it
|
||||
(troubleshooting §17).
|
||||
- eval_time: 90m
|
||||
alertname: FelisOTPAccountLocked
|
||||
exp_alerts: []
|
||||
- name: sign-in failure rate
|
||||
interval: 1m
|
||||
input_series:
|
||||
# Two doors failing at 3/min between them from t=0.
|
||||
- series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
|
||||
values: '0+2x40'
|
||||
- series: 'felis_auth_failures_total{door="op_login",reason="no_account",job="felis-api"}'
|
||||
values: '0+1x40'
|
||||
alert_rule_test:
|
||||
- eval_time: 8m
|
||||
alertname: FelisSignInFailures
|
||||
exp_alerts: []
|
||||
- eval_time: 25m
|
||||
alertname: FelisSignInFailures
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "over 30 refused sign-ins in 15 minutes"
|
||||
description: >-
|
||||
Wrong codes, unknown addresses or bad passkey assertions well above people
|
||||
mistyping: someone is guessing or enumerating. `sum by (door, reason)
|
||||
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
|
||||
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
|
||||
- name: people mistyping stays quiet
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_auth_failures_total{door="login_email",reason="bad_code",job="felis-api"}'
|
||||
values: '0 0 1 1 2 2 3 3 4 4 5x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 30m
|
||||
alertname: FelisSignInFailures
|
||||
exp_alerts: []
|
||||
- name: audit rows lost
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_audit_write_failures_total{job="felis-api",instance="api-0"}'
|
||||
values: '0 0 0 2x20'
|
||||
alert_rule_test:
|
||||
- eval_time: 2m
|
||||
alertname: FelisAuditWriteFailing
|
||||
exp_alerts: []
|
||||
- eval_time: 5m
|
||||
alertname: FelisAuditWriteFailing
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
job: felis-api
|
||||
instance: api-0
|
||||
exp_annotations:
|
||||
summary: "felis-api failed to write audit rows"
|
||||
description: >-
|
||||
The actions went through but their audit rows were lost. The felis-api log
|
||||
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
|
||||
being unreachable or out of disk.
|
||||
- name: operator and api presence
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_build_info{component="operator",version="v1",job="felis-operator",instance="op-0"}'
|
||||
values: '1x30'
|
||||
# felis-api stops being scraped after 5m; the series goes stale 5m later.
|
||||
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
|
||||
values: '1x5'
|
||||
alert_rule_test:
|
||||
- eval_time: 25m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisAPIDown
|
||||
exp_alerts: []
|
||||
- eval_time: 25m
|
||||
alertname: FelisAPIDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
component: api
|
||||
exp_annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- name: operator never scraped
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_build_info{component="api",version="v1",job="felis-api",instance="api-0"}'
|
||||
values: '1x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 5m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisOperatorDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
component: operator
|
||||
exp_annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- name: system and user servers down
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'felis_server_phase{server="login",role="login",phase="Starting",desired="Running"}'
|
||||
values: '1x20'
|
||||
- series: 'felis_server_phase{server="lobby",role="lobby",phase="Failed",desired="Running"}'
|
||||
values: '1x20'
|
||||
# Stopped on purpose: not an outage.
|
||||
- series: 'felis_server_phase{server="lobby2",role="lobby",phase="Stopped",desired="Stopped"}'
|
||||
values: '1x20'
|
||||
# A user server carries no role label (the operator publishes role="").
|
||||
- series: 'felis_server_phase{server="survival",phase="Failed",desired="Running"}'
|
||||
values: '1x20'
|
||||
- series: 'felis_server_phase{server="creative",phase="Running",desired="Running"}'
|
||||
values: '1x20'
|
||||
alert_rule_test:
|
||||
- eval_time: 9m
|
||||
alertname: FelisLoginGateDown
|
||||
exp_alerts: []
|
||||
- eval_time: 11m
|
||||
alertname: FelisLoginGateDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
server: login
|
||||
role: login
|
||||
phase: Starting
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "the login gate login is Starting"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver login`
|
||||
(troubleshooting §1, §2).
|
||||
- eval_time: 11m
|
||||
alertname: FelisSystemServerDown
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
server: lobby
|
||||
role: lobby
|
||||
phase: Failed
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "system server lobby (lobby) is Failed"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver lobby`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- eval_time: 3m
|
||||
alertname: FelisServerFailed
|
||||
exp_alerts: []
|
||||
- eval_time: 6m
|
||||
alertname: FelisServerFailed
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
server: survival
|
||||
phase: Failed
|
||||
desired: Running
|
||||
exp_annotations:
|
||||
summary: "server survival is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver survival`
|
||||
(troubleshooting §2).
|
||||
- name: operator reconcile errors and a stuck pass
|
||||
interval: 1m
|
||||
input_series:
|
||||
# Two failed reconciles a minute from 6m on.
|
||||
- series: 'controller_runtime_reconcile_errors_total{controller="minecraftserver",job="felis-operator"}'
|
||||
values: '0x5 0+2x20'
|
||||
# One pass that started at 5m and never returns.
|
||||
- series: 'workqueue_longest_running_processor_seconds{name="minecraftserver",controller="minecraftserver",job="felis-operator"}'
|
||||
values: '0x5 60+60x20'
|
||||
alert_rule_test:
|
||||
- eval_time: 5m
|
||||
alertname: FelisReconcileErrors
|
||||
exp_alerts: []
|
||||
- eval_time: 25m
|
||||
alertname: FelisReconcileErrors
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
exp_annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- eval_time: 12m
|
||||
alertname: FelisReconcileStuck
|
||||
exp_alerts: []
|
||||
- eval_time: 20m
|
||||
alertname: FelisReconcileStuck
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: critical
|
||||
exp_annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: world job failures and a late reaper
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="true"}'
|
||||
values: '0x2 1x10'
|
||||
- series: 'kube_job_failed{namespace="minecraft",job_name="backup-survival-abc",condition="false"}'
|
||||
values: '1x2 0x10'
|
||||
# A build Job in another namespace is FelisImageBuildFailures' business.
|
||||
- series: 'kube_job_failed{namespace="felis-build",job_name="build-x",condition="true"}'
|
||||
values: '1x12'
|
||||
# Evaluation starts at the epoch, so "over a day ago" is a negative timestamp.
|
||||
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
|
||||
values: '-100000x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 1m
|
||||
alertname: FelisWorldJobFailed
|
||||
exp_alerts: []
|
||||
- eval_time: 5m
|
||||
alertname: FelisWorldJobFailed
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
namespace: minecraft
|
||||
job_name: backup-survival-abc
|
||||
condition: "true"
|
||||
exp_annotations:
|
||||
summary: "Job backup-survival-abc failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/backup-survival-abc` has the error
|
||||
(troubleshooting §10).
|
||||
- eval_time: 5m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts: []
|
||||
- eval_time: 15m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts:
|
||||
- exp_labels:
|
||||
severity: warning
|
||||
namespace: minecraft
|
||||
cronjob: felis-reaper
|
||||
exp_annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: a reaper that ran yesterday stays quiet
|
||||
interval: 1m
|
||||
input_series:
|
||||
- series: 'kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"}'
|
||||
values: '-50000x30'
|
||||
alert_rule_test:
|
||||
- eval_time: 25m
|
||||
alertname: FelisReaperStale
|
||||
exp_alerts: []
|
||||
@@ -0,0 +1,278 @@
|
||||
# prometheus-operator twin of felis-alerts.yaml (kube-prometheus-stack loads
|
||||
# rules through the PrometheusRule CRD, not rule_files). The plain file is the
|
||||
# promtool-tested source; keep the groups in sync.
|
||||
apiVersion: monitoring.coreos.com/v1
|
||||
kind: PrometheusRule
|
||||
metadata:
|
||||
name: felis-alerts
|
||||
namespace: monitoring
|
||||
labels:
|
||||
# Change to match your stack's ruleSelector (kube-prometheus-stack's
|
||||
# default selects on the Helm release name).
|
||||
release: kube-prometheus-stack
|
||||
spec:
|
||||
groups:
|
||||
- name: felis.rules
|
||||
rules:
|
||||
- alert: FelisImageBuildFailures
|
||||
expr: increase(felis_image_build_failures_total[6h]) > 0
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "modpack/image build failed in the last 6h"
|
||||
description: >-
|
||||
felis_image_build_failures_total increased. Inspect the failed build Job
|
||||
(kubectl logs -n felis-build job/<build-job>); the same error text is on
|
||||
GET /api/v1/images/build/{id} and in the submitter's row in the panel.
|
||||
- alert: FelisSlowServerStarts
|
||||
expr: histogram_quantile(0.9, sum by (le) (rate(felis_start_duration_seconds_bucket[30m]))) > 300
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "p90 server start time exceeds 5 minutes"
|
||||
description: >-
|
||||
Starts regularly take over five minutes (felis_start_duration_seconds,
|
||||
observed when readiness is first reached). A start that never completes
|
||||
records nothing — cross-check desiredState=Running servers with no ready
|
||||
phase (troubleshooting §1).
|
||||
- name: felis.platform.rules
|
||||
rules:
|
||||
- alert: FelisOperatorDown
|
||||
expr: absent(felis_build_info{component="operator"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-operator is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="operator"} series for 10 minutes. Without
|
||||
the operator no server starts, stops or recovers. Check
|
||||
`kubectl -n felis get deploy felis-operator` and its log; if the pod is
|
||||
healthy, the felis-operator-metrics Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisAPIDown
|
||||
expr: absent(felis_build_info{component="api"})
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "felis-api is down or not scraped"
|
||||
description: >-
|
||||
No felis_build_info{component="api"} series for 10 minutes. The panel,
|
||||
sign-in and the proxy's player lookups all go through felis-api. Check
|
||||
`kubectl -n felis get deploy felis-api` and its log; if the pod is
|
||||
healthy, the felis-api-internal Service is not being scraped
|
||||
(troubleshooting §14).
|
||||
- alert: FelisLoginGateDown
|
||||
expr: felis_server_phase{role="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "the login gate {{ $labels.server }} is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Every player connection passes through the login server first, so no one
|
||||
can join. The MinecraftServer's conditions carry the reason:
|
||||
`kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §1, §2).
|
||||
- alert: FelisSystemServerDown
|
||||
expr: felis_server_phase{role!="",role!="login",desired="Running",phase!="Running"} == 1
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "system server {{ $labels.server }} ({{ $labels.role }}) is {{ $labels.phase }}"
|
||||
description: >-
|
||||
Players who sign in are sent to the lobby; while it is down they stay at
|
||||
the gate. `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
shows the reason (troubleshooting §1, §2).
|
||||
- alert: FelisServerFailed
|
||||
expr: felis_server_phase{role="",phase="Failed"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "server {{ $labels.server }} is Failed"
|
||||
description: >-
|
||||
The operator gave up on this server (a crash loop, an image that will not
|
||||
pull, a world volume that will not mount). Its conditions carry the
|
||||
reason: `kubectl -n minecraft describe minecraftserver {{ $labels.server }}`
|
||||
(troubleshooting §2).
|
||||
- alert: FelisReconcileErrors
|
||||
expr: sum(increase(controller_runtime_reconcile_errors_total{controller="minecraftserver"}[15m])) > 10
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the operator failed over 10 reconciles in 15 minutes"
|
||||
description: >-
|
||||
Server changes are being retried instead of applied. The felis-operator
|
||||
log names each failing server and its error.
|
||||
- alert: FelisReconcileStuck
|
||||
expr: max(workqueue_longest_running_processor_seconds{name="minecraftserver"}) > 300
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "an operator reconcile has been running for over 5 minutes"
|
||||
description: >-
|
||||
Each reconcile is bounded at 3 minutes, so this one is ignoring its
|
||||
deadline and holding a worker. The liveness probe restarts the operator
|
||||
once a pass passes 10 minutes; the log from before the restart shows
|
||||
where it hung.
|
||||
- name: felis.jobs.rules
|
||||
rules:
|
||||
- alert: FelisWorldJobFailed
|
||||
expr: kube_job_failed{namespace="minecraft",condition="true"} == 1
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "Job {{ $labels.job_name }} failed"
|
||||
description: >-
|
||||
A world backup, restore or reaper run failed; after a failed backup that
|
||||
world's newest archive is older than planned.
|
||||
`kubectl -n minecraft logs job/{{ $labels.job_name }}` has the error
|
||||
(troubleshooting §10).
|
||||
- alert: FelisReaperStale
|
||||
expr: time() - kube_cronjob_status_last_successful_time{namespace="minecraft",cronjob="felis-reaper"} > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the world reaper has not succeeded in over 26h"
|
||||
description: >-
|
||||
felis-reaper runs daily; idle worlds are neither backed up nor reclaimed
|
||||
while it fails. `kubectl -n minecraft get jobs --sort-by=.metadata.creationTimestamp`
|
||||
lists its runs, and the newest one's log
|
||||
shows why (troubleshooting §10).
|
||||
- name: felis.node.rules
|
||||
rules:
|
||||
- alert: FelisNodeDiskSpaceLow
|
||||
expr: >-
|
||||
node_filesystem_avail_bytes{fstype=~"ext4|xfs|btrfs"}
|
||||
/ node_filesystem_size_bytes{fstype=~"ext4|xfs|btrfs"} < 0.15
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "node filesystem {{ $labels.mountpoint }} below 15% available"
|
||||
description: >-
|
||||
Sustained disk pressure evicts game pods and garbage-collects images
|
||||
(troubleshooting §13b). Free space before kubelet raises DiskPressure.
|
||||
- alert: FelisNodeDiskPressure
|
||||
expr: kube_node_status_condition{condition="DiskPressure",status="true"} == 1
|
||||
for: 5m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "kubelet reports DiskPressure on {{ $labels.node }}"
|
||||
description: >-
|
||||
The eviction chain is in progress: control-plane pods hold
|
||||
system-cluster-critical and survive, game pods do not. Free disk now
|
||||
(troubleshooting §13b).
|
||||
- alert: FelisNodeMemoryLow
|
||||
expr: node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes < 0.10
|
||||
for: 15m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "node memory available below 10% for 15m"
|
||||
description: >-
|
||||
PostgreSQL, the control plane, the registry and game servers share one
|
||||
node; sustained memory pressure risks OOM kills.
|
||||
- name: felis.backup.rules
|
||||
rules:
|
||||
- alert: FelisDBBackupStale
|
||||
expr: time() - max(felis_db_backup_last_success_timestamp_seconds) > 26 * 3600
|
||||
for: 10m
|
||||
labels:
|
||||
severity: critical
|
||||
annotations:
|
||||
summary: "no control-plane database backup in over 26h"
|
||||
description: >-
|
||||
felis-db-backup.timer runs daily; the newest bundle is more than a day
|
||||
old. Read `journalctl -u felis-db-backup` on the host, then take one now
|
||||
with `sudo felis db backup` (troubleshooting §16).
|
||||
- alert: FelisDBBackupMetricMissing
|
||||
expr: absent(felis_db_backup_last_success_timestamp_seconds)
|
||||
for: 2h
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "database backup freshness is not being scraped"
|
||||
description: >-
|
||||
No felis_db_backup_last_success_timestamp_seconds series, so
|
||||
FelisDBBackupStale cannot fire. Point node-exporter's
|
||||
--collector.textfile.directory at the directory of
|
||||
FELIS_DB_BACKUP_METRICS (default /var/lib/node_exporter/textfile_collector)
|
||||
(troubleshooting §16).
|
||||
- name: felis.auth.rules
|
||||
rules:
|
||||
- alert: FelisMailBudgetExhausted
|
||||
expr: sum(increase(felis_mail_total{result="throttled"}[15m])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the install-wide mail budget refused mail"
|
||||
description: >-
|
||||
felis_mail_total{result="throttled"} increased: [smtp] max_per_hour is
|
||||
spent, and every sign-in code is refused with 429 mail_rate_limited until
|
||||
it refills. Check felis_rate_limited_total for a flood before raising the
|
||||
budget (troubleshooting §17).
|
||||
- alert: FelisMailDeliveryFailing
|
||||
expr: sum(increase(felis_mail_total{result="failed"}[15m])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "the SMTP relay refused mail in the last 15m"
|
||||
description: >-
|
||||
felis_mail_total{result="failed"} increased: sign-in codes are not being
|
||||
delivered (502 mail_undeliverable). The relay's reason is in the
|
||||
felis-api log (troubleshooting §17).
|
||||
- alert: FelisSignInFlood
|
||||
expr: sum(rate(felis_rate_limited_total{scope="auth_door"}[5m])) * 60 > 10
|
||||
for: 10m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "sign-in doors refusing over 10 requests a minute"
|
||||
description: >-
|
||||
The per-address sign-in limit has been refusing callers for 10 minutes.
|
||||
A script is hammering the auth doors; if real users report rate_limited
|
||||
at once instead, [auth] client_ip_header is missing and everyone shares
|
||||
the proxy's address (troubleshooting §17).
|
||||
- alert: FelisOTPAccountLocked
|
||||
expr: sum by (purpose) (increase(felis_auth_otp_lockouts_total[1h])) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "an account's email-code sign-in locked after 10 wrong codes"
|
||||
description: >-
|
||||
Someone entered 10 wrong codes for one account within 24h ({{ $labels.purpose }}).
|
||||
The audit log names the account (action auth.otp.locked); the owner was
|
||||
mailed. Unless they fumbled codes, someone is guessing at it
|
||||
(troubleshooting §17).
|
||||
- alert: FelisSignInFailures
|
||||
expr: sum(increase(felis_auth_failures_total[15m])) > 30
|
||||
for: 5m
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "over 30 refused sign-ins in 15 minutes"
|
||||
description: >-
|
||||
Wrong codes, unknown addresses or bad passkey assertions well above people
|
||||
mistyping: someone is guessing or enumerating. `sum by (door, reason)
|
||||
(increase(felis_auth_failures_total[15m]))` shows where; the audit rows
|
||||
(action auth.<door>.failed) carry each caller's client_ip (troubleshooting §17).
|
||||
- alert: FelisAuditWriteFailing
|
||||
expr: increase(felis_audit_write_failures_total[15m]) > 0
|
||||
labels:
|
||||
severity: warning
|
||||
annotations:
|
||||
summary: "felis-api failed to write audit rows"
|
||||
description: >-
|
||||
The actions went through but their audit rows were lost. The felis-api log
|
||||
names each lost row (`audit: lost ...`); the usual cause is PostgreSQL
|
||||
being unreachable or out of disk.
|
||||
+1733
-173
File diff suppressed because it is too large.
Load diff
File diff suppressed because it is too large.
Load diff
@@ -123,10 +123,6 @@ spec:
|
||||
lifecycle:
|
||||
description: Lifecycle tunes graceful shutdown (spec §7).
|
||||
properties:
|
||||
preStopSaveAndStop:
|
||||
description: PreStopSaveAndStop enables the operator-injected
|
||||
RCON save+stop preStop.
|
||||
type: boolean
|
||||
terminationGracePeriodSeconds:
|
||||
description: TerminationGracePeriodSeconds is the pod grace period
|
||||
(default 300).
|
||||
@@ -347,6 +343,14 @@ spec:
|
||||
- type
|
||||
type: object
|
||||
type: array
|
||||
emptySince:
|
||||
description: |-
|
||||
EmptySince is when the operator first observed 0 online players during
|
||||
a Running phase (spec §8 idle auto-stop). It is reset when a player joins
|
||||
or the server stops, so the empty-duration counter starts fresh each time
|
||||
the server becomes unoccupied.
|
||||
format: date-time
|
||||
type: string
|
||||
endpoint:
|
||||
description: Endpoint is where the proxy should route traffic.
|
||||
properties:
|
||||
|
||||
+25
-74
@@ -1,29 +1,28 @@
|
||||
#!/bin/bash
|
||||
# demo-up.sh — one-shot Felis demo bring-up.
|
||||
#
|
||||
# Collapses the four manual steps (bootstrap -> build/import limbo+lobby images ->
|
||||
# edit felis.toml -> felis setup) into a single command:
|
||||
#
|
||||
# sudo bash deploy/demo-up.sh
|
||||
#
|
||||
# It ends by exec'ing the interactive `felis setup` TUI (create the Owner account) —
|
||||
# that human step is the only thing this script cannot do for you.
|
||||
# Every piece a demo box needs — base platform (k3s + felis + docker + cloudflared +
|
||||
# control plane), the limbo/lobby/paper images, the felis-velocity proxy plugin, and
|
||||
# the [velocity] wiring in felis.host.toml — is built by deploy/bootstrap.sh. This
|
||||
# wrapper adds only the one step the installer cannot do: the interactive
|
||||
# `felis setup` TUI that creates the Owner account.
|
||||
#
|
||||
# Image source, in order of preference:
|
||||
# 1. Prebuilt tars at deploy/images/felis-limbo.tar + felis-lobby.tar (imported as-is).
|
||||
# 2. Otherwise built on this host with docker, resolving the LOOHP/Limbo CI jar and
|
||||
# the latest stable Paper jar automatically. Override any of:
|
||||
# LIMBO_JAR_URL LIMBO_SCHEM_URL LIMBO_VERSION PAPER_JAR_URL PAPER_MC_VERSION
|
||||
# It used to rebuild the game images here with its own copy of that logic, written
|
||||
# before bootstrap grew the job. The copy drifted: it pinned Paper 1.21.8 while the
|
||||
# installer derives one version from the Limbo login gate (both hops of a login must
|
||||
# speak one protocol), never built felis-velocity.jar (so the proxy it wired had
|
||||
# nowhere to route), and left docker running. The installer is the single origin.
|
||||
#
|
||||
# Toggles: SKIP_BOOTSTRAP=1 (base already up), SKIP_SETUP=1 (stop before the TUI).
|
||||
# Toggles: SKIP_BOOTSTRAP=1 (base + game stack already installed by a full bootstrap),
|
||||
# SKIP_SETUP=1 (stop before the TUI).
|
||||
set -Eeuo pipefail
|
||||
|
||||
SRC_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
|
||||
STATE_DIR=/etc/felis
|
||||
HOST_TOML="$STATE_DIR/felis.host.toml"
|
||||
IMG_DIR="$SRC_DIR/deploy/images"
|
||||
LIMBO_IMAGE="felis-limbo:demo"
|
||||
LOBBY_IMAGE="felis-lobby:demo"
|
||||
PLUGIN_JAR=/opt/felis/velocity/plugins/felis-velocity.jar
|
||||
K3S=/usr/local/bin/k3s
|
||||
FELIS=/usr/local/bin/felis
|
||||
|
||||
@@ -32,11 +31,11 @@ die() { printf '\033[1;31mERROR: %s\033[0m\n' "$*" >&2; exit 1; }
|
||||
|
||||
[ "$(id -u)" -eq 0 ] || die "run as root (sudo bash deploy/demo-up.sh)"
|
||||
|
||||
# 1. base platform (k3s + felis + docker + cloudflared + control plane) ----------
|
||||
# 1. base platform + full game stack (deploy/bootstrap.sh) -----------------------
|
||||
if [ "${SKIP_BOOTSTRAP:-0}" = 1 ]; then
|
||||
log "SKIP_BOOTSTRAP=1 — assuming the base platform is already up"
|
||||
else
|
||||
log "bringing up the base platform (deploy/bootstrap.sh)"
|
||||
log "bringing up the base platform and the game stack (deploy/bootstrap.sh)"
|
||||
bash "$SRC_DIR/deploy/bootstrap.sh"
|
||||
fi
|
||||
|
||||
@@ -45,67 +44,19 @@ command -v "$K3S" >/dev/null 2>&1 || K3S=k3s
|
||||
command -v "$K3S" >/dev/null 2>&1 || die "k3s not found — did bootstrap complete?"
|
||||
command -v "$FELIS" >/dev/null 2>&1 || die "felis not found — did bootstrap complete?"
|
||||
|
||||
# 2. get the two game images into k3s containerd --------------------------------
|
||||
if [ -f "$IMG_DIR/felis-limbo.tar" ] && [ -f "$IMG_DIR/felis-lobby.tar" ]; then
|
||||
log "importing prebuilt image tars from $IMG_DIR"
|
||||
"$K3S" ctr images import "$IMG_DIR/felis-limbo.tar"
|
||||
"$K3S" ctr images import "$IMG_DIR/felis-lobby.tar"
|
||||
# Optional: the plain-Paper recommended base, if a tar was staged for it.
|
||||
[ -f "$IMG_DIR/felis-paper.tar" ] && "$K3S" ctr images import "$IMG_DIR/felis-paper.tar"
|
||||
else
|
||||
log "no prebuilt tars in $IMG_DIR — building on this host with docker"
|
||||
command -v docker >/dev/null 2>&1 || die "docker not found; cannot build images"
|
||||
|
||||
rel=$(curl -fsSL --max-time 30 "https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/api/json" \
|
||||
| grep -oE 'target/Limbo-[0-9][^"]+\.jar' | head -1) || true
|
||||
: "${LIMBO_JAR_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/$rel}"
|
||||
: "${LIMBO_SCHEM_URL:=https://ci.loohpjames.com/job/Limbo/lastSuccessfulBuild/artifact/spawn.schem}"
|
||||
: "${LIMBO_VERSION:=$(basename "$rel" | sed -E 's/^Limbo-//; s/\.jar$//; s/-[0-9]+\.[0-9]+$//')}"
|
||||
[ -n "$rel" ] || [ -n "${LIMBO_JAR_URL##*artifact/}" ] || die "could not resolve the Limbo jar; set LIMBO_JAR_URL"
|
||||
log "building $LIMBO_IMAGE (Limbo $LIMBO_VERSION)"
|
||||
docker build -f "$SRC_DIR/deploy/limbo/Dockerfile" \
|
||||
--build-arg LIMBO_JAR_URL="$LIMBO_JAR_URL" \
|
||||
--build-arg LIMBO_SCHEM_URL="$LIMBO_SCHEM_URL" \
|
||||
--build-arg LIMBO_VERSION="$LIMBO_VERSION" \
|
||||
-t "$LIMBO_IMAGE" "$SRC_DIR"
|
||||
docker save "$LIMBO_IMAGE" | "$K3S" ctr images import -
|
||||
|
||||
: "${PAPER_MC_VERSION:=1.21.8}"
|
||||
: "${PAPER_JAR_URL:=$(curl -fsSL --max-time 30 "https://fill.papermc.io/v3/projects/paper/versions/${PAPER_MC_VERSION}/builds/latest" | grep -oE 'https://fill-data\.papermc\.io/[^"]+\.jar' | head -1)}"
|
||||
[ -n "$PAPER_JAR_URL" ] || die "could not resolve the Paper jar; set PAPER_JAR_URL"
|
||||
log "building $LOBBY_IMAGE (Paper $PAPER_MC_VERSION)"
|
||||
docker build -f "$SRC_DIR/deploy/lobby/Dockerfile" \
|
||||
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
|
||||
-t "$LOBBY_IMAGE" "$SRC_DIR"
|
||||
docker save "$LOBBY_IMAGE" | "$K3S" ctr images import -
|
||||
|
||||
# Plain Paper recommended base — same PAPER_JAR_URL, no plugins, no secret gate.
|
||||
: "${PAPER_IMAGE:=felis-paper:demo}"
|
||||
log "building $PAPER_IMAGE (plain Paper $PAPER_MC_VERSION, forwarding via the operator initContainer)"
|
||||
docker build -f "$SRC_DIR/deploy/paper/Dockerfile" \
|
||||
--build-arg PAPER_JAR_URL="$PAPER_JAR_URL" \
|
||||
-t "$PAPER_IMAGE" "$SRC_DIR"
|
||||
docker save "$PAPER_IMAGE" | "$K3S" ctr images import -
|
||||
fi
|
||||
|
||||
# 3. wire the images into the config `felis setup` reads ------------------------
|
||||
log "wiring [velocity] images into $HOST_TOML"
|
||||
# SKIP_BOOTSTRAP=1 trusts an earlier run to be complete. Check that it actually left
|
||||
# the full stack behind: a base from before the game-stack installer, or one whose
|
||||
# pieces were pruned by hand, must fail here with a pointer — not present as a proxy
|
||||
# that accepts logins and routes nowhere, with nothing in any log to say why.
|
||||
[ -f "$HOST_TOML" ] || die "missing $HOST_TOML — did bootstrap run?"
|
||||
if grep -q '^\[velocity\]' "$HOST_TOML"; then
|
||||
echo " [velocity] table already present — leaving it untouched"
|
||||
else
|
||||
cat >> "$HOST_TOML" <<EOF
|
||||
grep -q '^\[velocity\]' "$HOST_TOML" \
|
||||
|| die "$HOST_TOML has no [velocity] section — re-run the installer without SKIP_BOOTSTRAP so the system servers get wired"
|
||||
[ -f "$PLUGIN_JAR" ] \
|
||||
|| die "$PLUGIN_JAR missing — this base did not finish the full installer, and a proxy without it silently routes nothing; re-run the installer without SKIP_BOOTSTRAP"
|
||||
|
||||
[velocity]
|
||||
login_image = "$LIMBO_IMAGE"
|
||||
lobby_image = "$LOBBY_IMAGE"
|
||||
EOF
|
||||
echo " appended login_image=$LIMBO_IMAGE / lobby_image=$LOBBY_IMAGE"
|
||||
fi
|
||||
|
||||
# 4. interactive Owner creation + system-server provisioning -------------------
|
||||
# 2. interactive Owner creation + system-server provisioning --------------------
|
||||
if [ "${SKIP_SETUP:-0}" = 1 ]; then
|
||||
log "SKIP_SETUP=1 — base + images + config ready. Finish with: sudo felis setup"
|
||||
log "SKIP_SETUP=1 — base + game stack ready. Finish with: sudo felis setup"
|
||||
else
|
||||
log "launching 'felis setup' — create the Owner account (this is the only interactive step)"
|
||||
exec "$FELIS" setup
|
||||
|
||||
@@ -0,0 +1,23 @@
|
||||
# The upstream builds this release installs. deploy/bootstrap.sh reads this file (the
|
||||
# default FELIS_GAME_STACK=pinned), downloads exactly these artifacts and refuses any whose
|
||||
# sha256 differs, so every host installing one release gets the same login gate, lobby,
|
||||
# plain-Paper image and proxy, and a rerun rebuilds nothing that did not change.
|
||||
#
|
||||
# MC_VERSION is the protocol the whole stack speaks: Limbo speaks exactly one, and Paper
|
||||
# follows it so a client that passes the login gate can also reach the lobby.
|
||||
#
|
||||
# Refresh with deploy/update-game-stack-lock.sh, which resolves upstream's newest builds and
|
||||
# hashes them. Plain KEY=value lines only; bootstrap reads it without evaluating it.
|
||||
MC_VERSION=26.3
|
||||
LIMBO_VERSION=2026.0.3-ALPHA
|
||||
LIMBO_JAR_URL=https://ci.loohpjames.com/job/Limbo/76/artifact/target/Limbo-2026.0.3-ALPHA-26.3.jar
|
||||
LIMBO_JAR_SHA256=a2de91fcaa2255aed8a111b8786f370213423c7798eee778c00d016d83aa65d2
|
||||
LIMBO_SCHEM_URL=https://ci.loohpjames.com/job/Limbo/76/artifact/spawn.schem
|
||||
LIMBO_SCHEM_SHA256=70c85dae2db157971ef513e318820c5a5e10a96b813b70e21f8af2b632cbcbc7
|
||||
PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/49399919246cbf443efc8507447dc948eb7477c41be560b0e87e2a455aff824a/paper-26.3-40.jar
|
||||
PAPER_JAR_SHA256=49399919246cbf443efc8507447dc948eb7477c41be560b0e87e2a455aff824a
|
||||
LUCKPERMS_JAR_URL=https://download.luckperms.net/1672/bukkit/loader/LuckPerms-Bukkit-5.5.85.jar
|
||||
LUCKPERMS_JAR_SHA256=dc637ce18f48d3b75a7ffd1784b85be16a627359090dfe5adf15dab6d145dc7d
|
||||
VELOCITY_VERSION=3.5.1
|
||||
VELOCITY_JAR_URL=https://fill-data.papermc.io/v1/objects/b4e3164df5377346854dc6cb9e6a78022b1946ff69e89676313f5f6f1c6f0fb3/velocity-3.5.1-615.jar
|
||||
VELOCITY_JAR_SHA256=b4e3164df5377346854dc6cb9e6a78022b1946ff69e89676313f5f6f1c6f0fb3
|
||||
+27
-7
@@ -10,10 +10,11 @@
|
||||
# There is no bundled server.properties — Limbo writes a default on first run.
|
||||
# So the runtime is assembled from those two URLs (not a zip) via --build-arg:
|
||||
#
|
||||
# . <(grep '^LIMBO_' deploy/game-stack.lock)
|
||||
# docker build -f deploy/limbo/Dockerfile \
|
||||
# --build-arg LIMBO_JAR_URL=https://ci.loohpjames.com/job/Limbo/<n>/artifact/target/Limbo-<ver>.jar \
|
||||
# --build-arg LIMBO_SCHEM_URL=https://ci.loohpjames.com/job/Limbo/<n>/artifact/spawn.schem \
|
||||
# --build-arg LIMBO_VERSION=<maven-api-version> \
|
||||
# --build-arg LIMBO_JAR_URL="$LIMBO_JAR_URL" --build-arg LIMBO_JAR_SHA256="$LIMBO_JAR_SHA256" \
|
||||
# --build-arg LIMBO_SCHEM_URL="$LIMBO_SCHEM_URL" --build-arg LIMBO_SCHEM_SHA256="$LIMBO_SCHEM_SHA256" \
|
||||
# --build-arg LIMBO_VERSION="$LIMBO_VERSION" \
|
||||
# -t felis-limbo:demo .
|
||||
#
|
||||
# Note LIMBO_VERSION (the maven API version the plugin compiles against, e.g.
|
||||
@@ -33,7 +34,7 @@
|
||||
# JDK 17 fails to read them with "wrong version 65.0, should be 61.0". The image
|
||||
# also provides the `gradle` binary (this tree vendors no Gradle wrapper).
|
||||
# build.gradle still targets release 17 bytecode so the plugin loads on Java 17+.
|
||||
FROM gradle:8.14-jdk21 AS plugin
|
||||
FROM gradle:8.14-jdk21@sha256:5c4c0c4284de4a19951e82ac78f86dbcda2e136644bbfe159beba7ea3420cc80 AS plugin
|
||||
WORKDIR /src
|
||||
# Copy what the limbo module needs: its own tree plus the shared link core it
|
||||
# srcDir-includes (../shared/src/main/java → /src/plugins/shared/src/main/java), so
|
||||
@@ -50,21 +51,33 @@ RUN cd plugins/limbo \
|
||||
# 21-jre: the Limbo jar is Java 21 bytecode (class-file major 65), so a Java 17
|
||||
# JRE cannot run it (UnsupportedClassVersionError). A 21 JRE also runs the
|
||||
# plugin's release-17 bytecode fine.
|
||||
FROM eclipse-temurin:21-jre
|
||||
FROM eclipse-temurin:21-jre@sha256:49e21e16e3c86eb7816a44a67549910ed090fbeb40c29c525d58bf5e02e91b0f
|
||||
ARG LIMBO_JAR_URL
|
||||
ARG LIMBO_JAR_SHA256
|
||||
ARG LIMBO_SCHEM_URL
|
||||
ARG LIMBO_SCHEM_SHA256
|
||||
WORKDIR /limbo
|
||||
# Pull the two loose LOOHP/Limbo CI artifacts: the server jar (required, saved as
|
||||
# Limbo.jar) and the default spawn schematic (optional). Fail loudly if the jar
|
||||
# URL was not supplied.
|
||||
# Limbo.jar) and the default spawn schematic (optional). Each is checked against the
|
||||
# digest deploy/game-stack.lock names (bootstrap.sh passes it): the login gate is the
|
||||
# first thing every player's connection reaches, and Limbo's CI publishes no digest of
|
||||
# its own.
|
||||
RUN set -eu; \
|
||||
if [ -z "${LIMBO_JAR_URL:-}" ]; then \
|
||||
echo "ERROR: --build-arg LIMBO_JAR_URL=<Limbo server jar> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -z "${LIMBO_JAR_SHA256:-}" ]; then \
|
||||
echo "ERROR: --build-arg LIMBO_JAR_SHA256=<Limbo jar sha256> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -n "${LIMBO_SCHEM_URL:-}" ] && [ -z "${LIMBO_SCHEM_SHA256:-}" ]; then \
|
||||
echo "ERROR: --build-arg LIMBO_SCHEM_SHA256=<spawn.schem sha256> is required with LIMBO_SCHEM_URL" >&2; exit 1; \
|
||||
fi; \
|
||||
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
|
||||
curl -fSL "$LIMBO_JAR_URL" -o /limbo/Limbo.jar; \
|
||||
echo "$LIMBO_JAR_SHA256 /limbo/Limbo.jar" | sha256sum -c; \
|
||||
if [ -n "${LIMBO_SCHEM_URL:-}" ]; then \
|
||||
curl -fSL "$LIMBO_SCHEM_URL" -o /limbo/spawn.schem; \
|
||||
echo "$LIMBO_SCHEM_SHA256 /limbo/spawn.schem" | sha256sum -c; \
|
||||
fi; \
|
||||
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
|
||||
mkdir -p /limbo/plugins
|
||||
@@ -78,6 +91,13 @@ COPY deploy/limbo/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
|
||||
# The operator mounts the world PVC at /data. Runtime state lives there; /limbo
|
||||
# remains the immutable image seed copied into the volume by the entrypoint.
|
||||
WORKDIR /data
|
||||
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
|
||||
# the pod securityContext whatever USER an image declares; declaring it here as well
|
||||
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
|
||||
# lets that run write its world. The jar seed above stays root-owned and read-only to
|
||||
# the server.
|
||||
RUN chown 1000:1000 /data
|
||||
USER 1000:1000
|
||||
|
||||
ENV FELIS_HEALTH_PORT=8080
|
||||
# FELIS_GAME_PORT is the port the entrypoint pins Limbo to; it MUST equal the operator's
|
||||
|
||||
+18
-8
@@ -90,14 +90,21 @@ docker build -f deploy/limbo/Dockerfile \
|
||||
version `2026.0.2-ALPHA` (the `-26.2` CI qualifier is not published to the
|
||||
maven repo).
|
||||
|
||||
Import into k3s and point config at it:
|
||||
Publish it into the cluster's registry and point config at it. On the node
|
||||
itself (docker treats `127.0.0.1` as insecure by default):
|
||||
|
||||
```
|
||||
docker save felis-limbo:demo | sudo k3s ctr images import -
|
||||
# felis.toml → [velocity] login_image = "felis-limbo:demo"
|
||||
docker tag felis-limbo:demo 127.0.0.1:5000/felis/limbo:demo
|
||||
docker push 127.0.0.1:5000/felis/limbo:demo
|
||||
# felis.toml → [velocity] login_image = "registry.felis.svc:5000/felis/limbo:demo"
|
||||
sudo felis setup
|
||||
```
|
||||
|
||||
The registry keys a repository by the path after the host, so pushing through a
|
||||
`kubectl -n felis port-forward svc/registry 5000:5000` from another machine is
|
||||
equivalent. Hosting the image in the registry (rather than only importing it
|
||||
into containerd) is what lets kubelet re-pull it after an image GC.
|
||||
|
||||
## Ports (handled for you)
|
||||
|
||||
The entrypoint (`deploy/limbo/entrypoint.sh`) pins Limbo's `server-port` to
|
||||
@@ -136,11 +143,14 @@ set them by hand:
|
||||
pod's internal port 8081. That Service is deliberately separate from the external
|
||||
NodePort `felis-api` (443) so the no-Zero-Trust internal face is never published on
|
||||
a node's external IP.
|
||||
- **NetworkPolicy:** none is required today — neither the minecraft-namespace egress
|
||||
nor the control-namespace ingress is policy-locked, so the login pod's call to the
|
||||
API internal port is reachable. If a future deployment adds a minecraft egress lock
|
||||
or a control-namespace ingress fence, it must also open the login-pod →
|
||||
felis-api-internal (8081) path.
|
||||
- **NetworkPolicy:** the minecraft namespace is egress-locked
|
||||
(`felis-server-egress`: DNS plus the public internet, every private range
|
||||
excluded), so the internal API is unreachable from a game server by default.
|
||||
`felis-login-to-internal-api` opens exactly the login pod → felis-api (8081) path,
|
||||
selecting on the reserved `login` name AND the setup-owned
|
||||
`felis.lolicon.best/system-role=login` label the operator copies onto the pod — the
|
||||
same pair that decides who receives `FELIS_SERVICE_TOKEN`, so a user server cannot
|
||||
match it by picking a name.
|
||||
|
||||
The Velocity gate/lobby wiring is printed by `felis setup` and enforces the
|
||||
invariant: fresh connections hit `login` first, and only an authenticated release
|
||||
|
||||
+32
-10
@@ -5,12 +5,13 @@
|
||||
# the POST-auth /menu hub: it is reached only when the login gate transfers an
|
||||
# authenticated player onward, and it must never be a fallback target.
|
||||
#
|
||||
# Build (deploy/bootstrap.sh does this for you; the PAPER_JAR_URL comes from PaperMC's
|
||||
# Fill v3 API — api.papermc.io v2 has returned HTTP 410 since 2026-07-01):
|
||||
# Build (deploy/bootstrap.sh does this for you, with the URLs and digests
|
||||
# deploy/game-stack.lock names):
|
||||
# . <(grep -E '^(PAPER|LUCKPERMS)_' deploy/game-stack.lock)
|
||||
# docker build -f deploy/lobby/Dockerfile \
|
||||
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-26.2-<build>.jar \
|
||||
# --build-arg LUCKPERMS_JAR_URL="$(curl -fsSL https://metadata.luckperms.net/data/all \
|
||||
# | grep -o 'https://download.luckperms.net/[^"]*/bukkit/loader/[^"]*\.jar')" \
|
||||
# --build-arg PAPER_JAR_URL="$PAPER_JAR_URL" --build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
|
||||
# --build-arg LUCKPERMS_JAR_URL="$LUCKPERMS_JAR_URL" \
|
||||
# --build-arg LUCKPERMS_JAR_SHA256="$LUCKPERMS_JAR_SHA256" \
|
||||
# -t felis-lobby:demo .
|
||||
# docker save felis-lobby:demo | sudo k3s ctr images import -
|
||||
# # felis.toml → [velocity] lobby_image = "felis-lobby:demo"
|
||||
@@ -27,7 +28,7 @@
|
||||
# ---- build the felis-paper plugin jar (Paper API is Java 21) ----
|
||||
# gradle:8.14-jdk21 — an official Gradle image on JDK 21 (this tree vendors no Gradle
|
||||
# wrapper, and a bare JDK image ships no `gradle`). JDK 21 matches the Paper API.
|
||||
FROM gradle:8.14-jdk21 AS plugin
|
||||
FROM gradle:8.14-jdk21@sha256:5c4c0c4284de4a19951e82ac78f86dbcda2e136644bbfe159beba7ea3420cc80 AS plugin
|
||||
WORKDIR /src
|
||||
COPY plugins/paper/ ./plugins/paper/
|
||||
COPY plugins/shared/ ./plugins/shared/
|
||||
@@ -40,28 +41,42 @@ RUN cd plugins/paper \
|
||||
# 25-jre, not 21: Paper 26.2 declares `java.version.minimum = 25` (PaperMC Fill v3,
|
||||
# GET /v3/projects/paper/versions/26.2) and refuses to boot on anything older. A 25 JRE
|
||||
# also runs the plugin's Java-21 bytecode, so only the runtime moves.
|
||||
FROM eclipse-temurin:25-jre
|
||||
FROM eclipse-temurin:25-jre@sha256:bb036ed6cfdc57e3da7c22634d15f1b840d2caf76183861c80e81ca4b5104abb
|
||||
ARG PAPER_JAR_URL
|
||||
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
|
||||
# that shape at build time. Checking the digest after the download turns a truncated or
|
||||
# tampered fetch into a failed build instead of a lobby booted on the wrong bytes.
|
||||
ARG PAPER_JAR_SHA256
|
||||
# LuckPerms is required, not optional: the panel's whole permission surface
|
||||
# (internal/api/handlers_access.go) issues `lp user ...` over RCON, so a lobby built
|
||||
# without it answers every grant with "Unknown command" — a failure the operator only
|
||||
# discovers in production, because the server itself starts and runs perfectly well.
|
||||
# Failing the build is the cheap place to notice. Resolved by URL rather than pinned
|
||||
# here for the same reason PAPER_JAR_URL is: bootstrap.sh asks upstream for the current
|
||||
# build, so this file does not go stale on every LuckPerms release.
|
||||
# Failing the build is the cheap place to notice. Passed in rather than pinned here for
|
||||
# the same reason PAPER_JAR_URL is: deploy/game-stack.lock names the build, so this file
|
||||
# does not change on every LuckPerms release. The digest is required like Paper's; the
|
||||
# jar runs inside the lobby with the server's full permissions.
|
||||
ARG LUCKPERMS_JAR_URL
|
||||
ARG LUCKPERMS_JAR_SHA256
|
||||
WORKDIR /paper
|
||||
RUN set -eu; \
|
||||
if [ -z "${PAPER_JAR_URL:-}" ]; then \
|
||||
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
|
||||
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -z "${LUCKPERMS_JAR_URL:-}" ]; then \
|
||||
echo "ERROR: --build-arg LUCKPERMS_JAR_URL=<luckperms bukkit jar> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -z "${LUCKPERMS_JAR_SHA256:-}" ]; then \
|
||||
echo "ERROR: --build-arg LUCKPERMS_JAR_SHA256=<luckperms jar sha256> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
|
||||
mkdir -p /paper/plugins; \
|
||||
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
|
||||
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
|
||||
curl -fSL "$LUCKPERMS_JAR_URL" -o /paper/plugins/LuckPerms.jar; \
|
||||
echo "$LUCKPERMS_JAR_SHA256 /paper/plugins/LuckPerms.jar" | sha256sum -c; \
|
||||
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*; \
|
||||
echo "eula=true" > /paper/eula.txt
|
||||
COPY --from=plugin /felis-paper.jar /paper/plugins/felis-paper.jar
|
||||
@@ -73,6 +88,13 @@ COPY deploy/lobby/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
|
||||
# The operator mounts the world PVC at /data. Runtime state lives there; /paper
|
||||
# remains the immutable image seed copied into the volume by the entrypoint.
|
||||
WORKDIR /data
|
||||
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
|
||||
# the pod securityContext whatever USER an image declares; declaring it here as well
|
||||
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
|
||||
# lets that run write its world. The jar seed above stays root-owned and read-only to
|
||||
# the server.
|
||||
RUN chown 1000:1000 /data
|
||||
USER 1000:1000
|
||||
|
||||
# FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's
|
||||
# GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep
|
||||
|
||||
@@ -29,9 +29,14 @@ this at every layer:
|
||||
```
|
||||
docker build -f deploy/lobby/Dockerfile \
|
||||
--build-arg PAPER_JAR_URL=https://<mirror>/paper-1.21.x-<build>.jar \
|
||||
--build-arg PAPER_JAR_SHA256=<sha256 of that jar> \
|
||||
-t felis-lobby:demo .
|
||||
docker save felis-lobby:demo | sudo k3s ctr images import -
|
||||
# felis.toml → [velocity] lobby_image = "felis-lobby:demo"
|
||||
# Publish into the cluster's registry (on the node; docker treats 127.0.0.1 as
|
||||
# insecure by default — or through a `kubectl -n felis port-forward svc/registry
|
||||
# 5000:5000`, which is equivalent: only the path after the host matters).
|
||||
docker tag felis-lobby:demo 127.0.0.1:5000/felis/lobby:demo
|
||||
docker push 127.0.0.1:5000/felis/lobby:demo
|
||||
# felis.toml → [velocity] lobby_image = "registry.felis.svc:5000/felis/lobby:demo"
|
||||
sudo felis setup
|
||||
```
|
||||
|
||||
|
||||
@@ -93,7 +93,7 @@ else
|
||||
echo " injects it from the <server>-rcon Secret when spec.rcon.enabled is true." >&2
|
||||
fi
|
||||
|
||||
# ponytail: rewritten whole, not merged. Paper loads this file and fills every key it does
|
||||
# Rewritten whole, not merged. Paper loads this file and fills every key it does
|
||||
# not find with the default, then writes the full tree back — so a proxies-only file is a
|
||||
# complete, stable input, and the lobby's other globals are simply always the defaults.
|
||||
# That is true of a system server Felis owns end to end; if admins are ever allowed to tune
|
||||
|
||||
+21
-5
@@ -5,7 +5,7 @@
|
||||
# internal/store/migrations/0019_recommended_paper.sql). It is NOT a system server: it
|
||||
# carries no felis-paper /menu plugin, no LuckPerms, and no forwarding-secret gate.
|
||||
#
|
||||
# It writes NO Velocity forwarding config itself. The operator injects a root
|
||||
# It writes NO Velocity forwarding config itself. The operator injects a
|
||||
# `felis init-forwarding` initContainer into every USER server (internal/operator/
|
||||
# builders.go: buildStatefulSet) that writes config/paper-global.yml + server.properties
|
||||
# online-mode=false onto the /data PVC before this container starts. That external step is
|
||||
@@ -14,10 +14,11 @@
|
||||
# image a user brings is made joinable the same way. If the initContainer is absent (no
|
||||
# FELIS_IMAGE configured) Paper boots as a standalone online server: degraded, not broken.
|
||||
#
|
||||
# Build (deploy/bootstrap.sh does this for you; PAPER_JAR_URL comes from PaperMC's Fill v3
|
||||
# API — the SAME url the lobby build resolves, so this reuses it and adds no new dependency):
|
||||
# Build (deploy/bootstrap.sh does this for you, with the SAME Paper build the lobby uses —
|
||||
# deploy/game-stack.lock names it — so this adds no new dependency):
|
||||
# . <(grep '^PAPER_' deploy/game-stack.lock)
|
||||
# docker build -f deploy/paper/Dockerfile \
|
||||
# --build-arg PAPER_JAR_URL=https://fill-data.papermc.io/v1/objects/<sha>/paper-<ver>-<build>.jar \
|
||||
# --build-arg PAPER_JAR_URL="$PAPER_JAR_URL" --build-arg PAPER_JAR_SHA256="$PAPER_JAR_SHA256" \
|
||||
# -t felis-paper:demo .
|
||||
# docker save felis-paper:demo | sudo k3s ctr images import -
|
||||
# # felis.toml → recommended via 0019_recommended_paper.sql (no [velocity] key points here)
|
||||
@@ -28,15 +29,23 @@
|
||||
|
||||
# 25-jre, not 21: Paper 26.2 declares java.version.minimum=25 (PaperMC Fill v3) and refuses
|
||||
# to boot on anything older.
|
||||
FROM eclipse-temurin:25-jre
|
||||
FROM eclipse-temurin:25-jre@sha256:bb036ed6cfdc57e3da7c22634d15f1b840d2caf76183861c80e81ca4b5104abb
|
||||
ARG PAPER_JAR_URL
|
||||
# Required alongside the URL: Fill's URLs are content-addressed, but nothing enforces
|
||||
# that shape at build time. Checking the digest after the download turns a truncated or
|
||||
# tampered fetch into a failed build instead of a server booted on the wrong bytes.
|
||||
ARG PAPER_JAR_SHA256
|
||||
RUN set -eu; \
|
||||
if [ -z "${PAPER_JAR_URL:-}" ]; then \
|
||||
echo "ERROR: --build-arg PAPER_JAR_URL=<paper jar> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
if [ -z "${PAPER_JAR_SHA256:-}" ]; then \
|
||||
echo "ERROR: --build-arg PAPER_JAR_SHA256=<paper jar sha256> is required" >&2; exit 1; \
|
||||
fi; \
|
||||
apt-get update && apt-get install -y --no-install-recommends curl ca-certificates; \
|
||||
mkdir -p /paper; \
|
||||
curl -fSL "$PAPER_JAR_URL" -o /paper/paper.jar; \
|
||||
echo "$PAPER_JAR_SHA256 /paper/paper.jar" | sha256sum -c; \
|
||||
apt-get purge -y curl && apt-get autoremove -y && rm -rf /var/lib/apt/lists/*
|
||||
COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
|
||||
|
||||
@@ -45,6 +54,13 @@ COPY deploy/paper/entrypoint.sh /usr/local/bin/felis-entrypoint.sh
|
||||
# on the PVC. /paper stays the immutable image seed: the jar is never copied onto the
|
||||
# volume, so the panel file editor (which sees only /data) cannot tamper with it.
|
||||
WORKDIR /data
|
||||
# Run as the game uid (naming.GameUID in the Go tree). The operator pins the same uid in
|
||||
# the pod securityContext whatever USER an image declares; declaring it here as well
|
||||
# keeps a plain `docker run` of this image off root, and chowning the empty /data seed
|
||||
# lets that run write its world. The jar seed above stays root-owned and read-only to
|
||||
# the server.
|
||||
RUN chown 1000:1000 /data
|
||||
USER 1000:1000
|
||||
|
||||
# FELIS_GAME_PORT is the port the entrypoint pins Paper to; it MUST equal the operator's
|
||||
# GamePort (internal/operator/builders.go). Default 25565 — override only in lockstep with
|
||||
|
||||
@@ -3,7 +3,7 @@
|
||||
#
|
||||
# A plain Paper backend for a user's OWN world — NOT a system server. Unlike deploy/limbo
|
||||
# and deploy/lobby it writes no Velocity forwarding config and has no secret gate: the
|
||||
# operator injects a root `felis init-forwarding` initContainer that writes
|
||||
# operator injects a `felis init-forwarding` initContainer that writes
|
||||
# config/paper-global.yml + server.properties online-mode=false onto /data BEFORE this
|
||||
# container starts, so forwarding is configured externally and this stays a drop-in Paper
|
||||
# image. With no initContainer (no FELIS_IMAGE) Paper just boots standalone-online —
|
||||
|
||||
Executable
+73
@@ -0,0 +1,73 @@
|
||||
#!/usr/bin/env bash
|
||||
# Refreshes deploy/game-stack.lock to upstream's newest builds, hashed.
|
||||
#
|
||||
# bash deploy/update-game-stack-lock.sh # rewrite the lock in place
|
||||
# bash deploy/update-game-stack-lock.sh --check # exit 1 if upstream moved on
|
||||
#
|
||||
# It runs the same resolver bootstrap.sh uses for FELIS_GAME_STACK=latest (the functions are
|
||||
# lifted out of bootstrap.sh, so the two cannot drift), then pins Velocity's newest build of
|
||||
# VELOCITY_LATEST_MINOR. Limbo and LuckPerms publish no digest, so their jars are downloaded
|
||||
# and hashed here; Paper and Velocity come from Fill's content-addressed URLs.
|
||||
#
|
||||
# Review the diff before committing: MC_VERSION moves the login gate's protocol, and the
|
||||
# lobby, plain-Paper image and every client follow it.
|
||||
set -Eeuo pipefail
|
||||
|
||||
here="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
BS="${here}/bootstrap.sh"
|
||||
LOCK="${here}/game-stack.lock"
|
||||
check=0
|
||||
case "${1:-}" in
|
||||
--check) check=1 ;;
|
||||
"") ;;
|
||||
*) printf 'usage: %s [--check]\n' "$0" >&2; exit 2 ;;
|
||||
esac
|
||||
|
||||
log() { printf '[lock] %s\n' "$*" >&2; }
|
||||
ok() { printf '[ ok ] %s\n' "$*" >&2; }
|
||||
warn() { :; }
|
||||
die() { printf '[fail] %s\n' "$*" >&2; exit 1; }
|
||||
|
||||
lift() { # function-name
|
||||
local body
|
||||
body="$(awk -v f="$1" '$0 ~ "^" f "\\(\\) \\{" {on=1} on {print} on && /^}/ {exit}' "$BS")"
|
||||
[ -n "$body" ] || die "bootstrap.sh no longer defines $1"
|
||||
eval "$body"
|
||||
}
|
||||
for fn in meta_get papermc_latest_jar luckperms_latest_jar url_sha256 resolve_latest_game_jars; do
|
||||
lift "$fn"
|
||||
done
|
||||
eval "$(grep '^VELOCITY_LATEST_MINOR=' "$BS")"
|
||||
[ -n "${VELOCITY_LATEST_MINOR:-}" ] || die "bootstrap.sh no longer sets VELOCITY_LATEST_MINOR"
|
||||
|
||||
resolve_latest_game_jars
|
||||
log "resolving the newest Velocity ${VELOCITY_LATEST_MINOR} build"
|
||||
velocity="$(papermc_latest_jar velocity "$VELOCITY_LATEST_MINOR")" \
|
||||
|| die "no Velocity build for ${VELOCITY_LATEST_MINOR}"
|
||||
# shellcheck disable=SC2034 # read back through ${!key} below
|
||||
VELOCITY_VERSION="$VELOCITY_LATEST_MINOR"
|
||||
# shellcheck disable=SC2034
|
||||
VELOCITY_JAR_URL="${velocity% *}"
|
||||
# shellcheck disable=SC2034
|
||||
VELOCITY_JAR_SHA256="${velocity##* }"
|
||||
|
||||
tmp="$(mktemp)"
|
||||
trap 'rm -f "$tmp"' EXIT
|
||||
# The comment header is kept as it is; only the KEY=value lines are regenerated.
|
||||
sed -n '/^#/p;/^#/!q' "$LOCK" > "$tmp"
|
||||
for key in MC_VERSION LIMBO_VERSION LIMBO_JAR_URL LIMBO_JAR_SHA256 LIMBO_SCHEM_URL LIMBO_SCHEM_SHA256 \
|
||||
PAPER_JAR_URL PAPER_JAR_SHA256 LUCKPERMS_JAR_URL LUCKPERMS_JAR_SHA256 \
|
||||
VELOCITY_VERSION VELOCITY_JAR_URL VELOCITY_JAR_SHA256; do
|
||||
printf '%s=%s\n' "$key" "${!key}" >> "$tmp"
|
||||
done
|
||||
|
||||
if cmp -s "$tmp" "$LOCK"; then
|
||||
ok "game-stack.lock already pins upstream's newest builds"
|
||||
exit 0
|
||||
fi
|
||||
diff -u "$LOCK" "$tmp" >&2 || true
|
||||
if [ "$check" = 1 ]; then
|
||||
die "upstream has newer builds than game-stack.lock"
|
||||
fi
|
||||
cp "$tmp" "$LOCK"
|
||||
ok "game-stack.lock updated; run go test . and the bootstrap tests, then commit"
|
||||
@@ -1,41 +0,0 @@
|
||||
# Foundational subsystems: the initial Felis import (ledger backfill)
|
||||
|
||||
- **Type:** feature (initial import) — retroactive ledger entry
|
||||
- **Date:** 2026-06-26
|
||||
- **Area:** `apis/`, `internal/` (naming, rcon, store, config, build, backup, operator,
|
||||
submit, api, platform), `cmd/felis`, `plugins/`
|
||||
- **Commits:**
|
||||
- `7fbebfe` feat(apis): MinecraftServer CRD types (v1alpha1) — the lifecycle source of truth (§1)
|
||||
- `708cdfc` feat(core): naming, RCON, store (Postgres + embedded migrations), config, image-build libraries
|
||||
- `43ab921` feat(backup): archive-based world backup/restore + the retention/idle reaper
|
||||
- `78b8cf6` feat(operator): MinecraftServer controller and reconcilers
|
||||
- `d39605e` feat(submit): user modpack build + admin-approval pipeline (see [modpack-submission-lane](2026-06-26-modpack-submission-lane.md))
|
||||
- `b508fcc` feat(api): dual-faced felis-api — permissions/LuckPerms, modpack lane, admin fleet read
|
||||
- `47fcd90` feat(platform): node orchestration + the `cmd/felis` single-binary entrypoint
|
||||
- `93f143f` feat(plugins): Velocity proxy + Fabric/Forge/NeoForge/Paper integration mods
|
||||
- **Tasks:** #23 (permissions), #24 (modpack lane), #25 (fleet read)
|
||||
|
||||
## What it did
|
||||
|
||||
Stood up the whole backend spine in one build-order sweep: the Kubernetes CRD that is
|
||||
the lifecycle source of truth, the core libraries (deterministic resource naming, the
|
||||
RCON client, the Postgres store with embedded SQL migrations, config loading, container
|
||||
image-build helpers), the backup/restore/reaper subsystems, the operator controller
|
||||
that drives `MinecraftServer` resources, the user-modpack submit+approval pipeline, the
|
||||
dual-faced (internal/external) felis-api behind a Zero-Trust guard, the platform
|
||||
orchestrator that wires it all together under `cmd/felis`, and the server-side
|
||||
integration plugins.
|
||||
|
||||
## Why
|
||||
|
||||
This is the project's first functional import — the substrate every later change edits.
|
||||
It predates the change-ledger convention (established `fad48ff`, 2026-07-06), so it never
|
||||
got a contemporaneous detail doc; this entry backfills one.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
|
||||
> change-ledger's detail-doc axis (§ Convention). This entry deliberately describes only
|
||||
> what these eight commits **introduced** on 2026-06-26 — the named subsystems have been
|
||||
> extended and reworked many times since (auth, passkey, metrics, quotas, updates), and
|
||||
> that later work lives in its own dated detail docs, not here. Not independently
|
||||
> re-verified for this doc; each subsystem was verified at its original commit and the
|
||||
> current tree builds green at `9911b8c` (WSL oracle, go1.26.4).
|
||||
@@ -1,30 +0,0 @@
|
||||
# Modpack submission lane: build/approval pipeline + storage backends (ledger backfill)
|
||||
|
||||
- **Type:** feature — retroactive ledger entry
|
||||
- **Date:** 2026-06-26 – 2026-07-02
|
||||
- **Area:** `internal/submit` (build/approval pipeline, storage backends), `internal/api` (submission endpoints)
|
||||
- **Commits:**
|
||||
- `d39605e` feat(submit): user modpack build + approval pipeline — an uploaded modpack stays `pending_review` and is never built until an admin approves; approval is a single-winner compare-and-swap handing off to the image-build Job, keeping the mandatory vulnerability scan in front of any push
|
||||
- `598f3d3` feat(submit): local + S3 backends for modpack upload contexts, installer-selectable
|
||||
- **Tasks:** #24 (§8 user-submitted modpack approval lane)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the user-directed extension over the image-build subsystem: a player uploads a
|
||||
modpack context, it sits in `pending_review`, and an admin's approval is the single-winner
|
||||
gate that hands off to the build Job — with the vulnerability scan always ahead of any
|
||||
registry push. `598f3d3` makes the upload-context store pluggable (local filesystem or S3),
|
||||
selectable at install time.
|
||||
|
||||
## Why
|
||||
|
||||
Untrusted user content must never build or push unreviewed, and the compare-and-swap
|
||||
approval guarantees exactly one build per submission even under a double-click or retry.
|
||||
The storage-backend choice lets a single-node demo use local disk while a real deployment
|
||||
uses S3, without a code change.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The approval
|
||||
> compare-and-swap and endpoints were unit-tested at their commits; the S3 path is
|
||||
> integration-configurable. The panel-side submission/approval UI is the collaborator's
|
||||
> frontend work and is tracked only by its INDEX rows. Not independently re-verified for
|
||||
> this doc; current tree green at `9911b8c`.
|
||||
@@ -1,38 +0,0 @@
|
||||
# Cloudflare Tunnel + Access edge (cfsetup) + NodePort fencing (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-07-01
|
||||
- **Area:** `internal/cfsetup` (pure core + integration runner), `cmd/felis` (TUI edge flow), edge nftables fence
|
||||
- **Commits:**
|
||||
- `53a7664` feat(cfsetup): recommended Cloudflare Tunnel + Access edge (§14) — domain- and IdP-agnostic; the load-bearing `validateFailClosed` allowlist refuses any policy that could be public; fail-shut 404 catch-all; the raw game host is never proxied
|
||||
- `ba13839` feat(breakglass): optional Tunnel + Access setup in the sudo TUI, an independent peer of Owner provisioning
|
||||
- `a531f5e` fix(cfsetup): keep the connector install in the host apply layer only (drop the duplicate `StartConnector`)
|
||||
- `2810fe8` fix(cfsetup): repoint a stale DNS record when routing a tunnel hostname
|
||||
- `7d3be64` feat(cfsetup): start the tunnel connector as a setup step
|
||||
- `346ec68` refactor(deploy): rework the cloudflare-edge walkthrough — restructured the edge TUI flow and added a tested `cfsetup` integration-runner path (with TUI height-measure/root tests)
|
||||
- `e058a64` feat(edge): close the panel NodePort to the public after the tunnel is up — nftables at prerouting `raw` (-300), before kube-proxy's NodePort DNAT, gated on the connector actually serving; loopback accepted first so the connector origin hop is untouched
|
||||
- **Tasks:** #37 (fence panel NodePort to public after tunnel)
|
||||
|
||||
## What it did
|
||||
|
||||
Stood up the optional one-click Zero-Trust edge: a Cloudflare Tunnel routing only the web
|
||||
hostnames plus a fail-closed Access application, provisioned from the sudo TUI against the
|
||||
operator's own Cloudflare account. `e058a64` then closes the Access-bypass hole where a
|
||||
direct `https://<node-ip>:<nodeport>/` with the right Host header reached the origin
|
||||
behind Access, by fencing the NodePort at the nftables raw hook so the packet is caught on
|
||||
its original destination port — but only once the connector is confirmed serving, so
|
||||
fencing never severs the only web path to a live origin.
|
||||
|
||||
## Why
|
||||
|
||||
Access is only a security boundary if the origin cannot be reached around it. The
|
||||
fail-closed policy guard (`validateFailClosed`) and the NodePort fence are the two
|
||||
load-bearing safety properties: a policy that could be public aborts the run with nothing
|
||||
created, and a routable-but-unfenced NodePort would defeat the whole edge.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The policy guard,
|
||||
> ingress generation, request bodies, gating, and the nftables ruleset shape / conn-count
|
||||
> gate are unit-tested; the live cloudflared/Cloudflare-API and `nft` calls are
|
||||
> INTEGRATION-ONLY (need a real account). KNOWN-LIMITATION: the fence targets nftables;
|
||||
> firewalld-native coordination is deferred. Not independently re-verified for this doc;
|
||||
> current tree green at `9911b8c`.
|
||||
@@ -1,32 +0,0 @@
|
||||
# Console auth: local-password login → passwordless migration (ledger backfill)
|
||||
|
||||
- **Type:** feature + refactor — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-07-04
|
||||
- **Area:** `internal/api` (auth handlers, sessions), `internal/store` (users schema)
|
||||
- **Commits:**
|
||||
- `af14f02` feat(api): local-password authentication backend — login/logout/change-password on `op.console`; HttpOnly+Secure+SameSite=Lax host-only server-side sessions (SHA-256, 12h TTL); anti-enumeration uniform bcrypt; JSON-only credential writes (415 otherwise); fails closed unless `local_auth_enabled`
|
||||
- `0c1cc59` feat(auth): migrate console login to passwordless
|
||||
- `3b43f05` refactor(api): drop the dead login concurrency limiter and reconcile passwordless comments
|
||||
- `c20b12c` refactor(api): drop the dead password-era `ResetMailer`, reconcile passkey-unbind docs
|
||||
- **Tasks:** #27 (B1 thin thread), #79/#80/#81 (residue sweep + primitive adjudication)
|
||||
|
||||
## What it did
|
||||
|
||||
Shipped the staff local-password door (`af14f02`) as the primary web login when
|
||||
Zero Trust is not in front of the API, then migrated the console to passwordless
|
||||
(`0c1cc59`) once email-OTP + passkey were the intended factors. The two refactors
|
||||
(`3b43f05`, `c20b12c`) then swept the password-era residue — the now-dead login
|
||||
concurrency limiter and the `ResetMailer` — so no unused password machinery lingered in
|
||||
the compile path, and reconciled the stale comments that referenced it.
|
||||
|
||||
## Why
|
||||
|
||||
`op.console` needs a real login even in deployments without a Cloudflare-Access edge; the
|
||||
password backend was that. Once the passwordless factors landed, keeping the old password
|
||||
scaffolding around was a bug farm — the sweep is the closeout evidence that the migration
|
||||
was complete, not half-done.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. `af14f02` was
|
||||
> covered by Go unit tests (content-type guard, anti-enumeration, forced-change lockdown)
|
||||
> at its commit. Not independently re-verified for this doc; current tree green at
|
||||
> `9911b8c` (WSL oracle, go1.26.4).
|
||||
@@ -1,38 +0,0 @@
|
||||
# Deploy: one-line bootstrap installer + demo bring-up (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-07-03
|
||||
- **Area:** `deploy/` (bootstrap.sh, Dockerfiles, demo-up.sh), image build context
|
||||
- **Commits:**
|
||||
- `58fa4b0` feat(deploy): one-line bootstrap installer + distroless felis image (auto-detects apt/dnf, installs Docker/k3s/PostgreSQL, opens pg_hba to the pod CIDR, runs migrations, applies the control-plane bundle, leaves Web disabled pending `felis setup`)
|
||||
- `94a3b7b` fix(deploy): harden bootstrap for RHEL-family Linux
|
||||
- `deaa2f8` feat(deploy): zypper support (openSUSE/SLES)
|
||||
- `318a724` feat(deploy): pacman support (Arch)
|
||||
- `e5f1682` refactor(deploy)!: TUI (breaking walkthrough restructure)
|
||||
- `28c3eee` refactor(deploy): improved TUI walkthrough
|
||||
- `c14ed17` fix(docker): keep embedded `panel/` and `deploy/` in the image build context
|
||||
- `d9e866f` fix(deploy): make the lobby image actually build (re-include `plugins/paper`, build on `gradle:8.14-jdk21`)
|
||||
- `b84debf` feat(deploy): one-shot `demo-up.sh` — bootstrap → build/import limbo+lobby images → wire `[velocity]` image refs → `felis setup`, ending in the interactive Owner TUI
|
||||
- **Tasks:** #26 (Phase A bootstrap verified end-to-end on the Demo VM)
|
||||
|
||||
## What it did
|
||||
|
||||
Made a bare Linux box a running Felis with one command. `bootstrap.sh` auto-detects the
|
||||
host package manager across the four major families (apt/dnf/zypper/pacman), installs
|
||||
whatever is missing (Docker, k3s, PostgreSQL, cloudflared), builds+imports the distroless
|
||||
felis image, opens `pg_hba` to the pod CIDR, runs migrations, and applies the rendered
|
||||
control-plane bundle. `demo-up.sh` wraps that plus the login-limbo/lobby image build and
|
||||
`felis setup` into a single command, stopping only at the Owner-creation TUI it cannot
|
||||
automate.
|
||||
|
||||
## Why
|
||||
|
||||
The spec calls for a self-hostable single-node deployment a SysAdmin can stand up without
|
||||
a Kubernetes background. The package-manager fan-out and the demo wrapper are what make
|
||||
"one line" true across real distros rather than only on the author's box.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history to close the
|
||||
> change-ledger's detail-doc axis. `deploy/` is shell + Dockerfiles (not Go-oracle
|
||||
> verifiable); `d9e866f` records a real build+boot check (limbo `/healthz` 200, lobby
|
||||
> reaches "Done"). Not independently re-verified for this doc; current tree green at
|
||||
> `9911b8c`.
|
||||
@@ -1,35 +0,0 @@
|
||||
# felis CLI: break-glass recovery console + first-run setup (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-06-30
|
||||
- **Area:** `cmd/felis` (break-glass/setup TUI, apply, migrate), `internal/api` (audit, owner store), `deploy/`
|
||||
- **Commits:**
|
||||
- `e108a37` feat(cli): break-glass emergency console TUI — root-only (`euid==0`), provisions/resets the Owner directly against Postgres, enables local login, prints a durable one-time-password summary
|
||||
- `2d0bbb0` feat(cli): attribute break-glass recovery to the SysAdmin who runs it — bootstrap / recovery (bcrypt) / root-override, each audited with an honest `verified` flag and payload
|
||||
- `a94b001` feat(deploy): break-glass Operator account provisioning
|
||||
- `eb5875a` feat(felis): Operator break-glass op behind an operation menu
|
||||
- `f5d00f3` feat(cli): `felis apply` for direct CRD creation
|
||||
- `9c46632` feat(cli): `felis setup` first-run console (shared `runConsoleTUI` model, reclaim protection, cfsetup idempotency, `[auth].admin_hostname` respect)
|
||||
- `7d91373` fix(migrate): honor `-config` placed after the `up` verb (flag.Parse stops at the first non-flag token)
|
||||
- **Tasks:** #27 (B1 login→change-pw→TUI reset)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the local-root recovery and first-run surface that bypasses web Zero Trust by
|
||||
design. `felis breakGlass` mints or resets the Owner when the web login is unreachable;
|
||||
`2d0bbb0` makes it accountable by recording *which* SysAdmin broke the glass across three
|
||||
audited modes. `felis setup` is the non-emergency first-run twin sharing the same console
|
||||
model. `felis apply` writes a `MinecraftServer` CRD directly, and `7d91373` fixes the
|
||||
`migrate` flag parse so a configured DB path after `up` is honored.
|
||||
|
||||
## Why
|
||||
|
||||
An operator with root on the node and a kubeconfig must always be able to recover the
|
||||
platform — that is break-glass's whole job, so it never refuses. Attribution
|
||||
(`2d0bbb0`) closes the gap that root is machine authority, not a human identity: the root
|
||||
gate is necessary but not sufficient for the audit trail.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The core logic was
|
||||
> covered by Go unit tests over a fake owner store at each commit (auth match/non-match,
|
||||
> the three audit modes, headless TUI drive). The bubbletea TUI glue is untested by house
|
||||
> convention. Not independently re-verified for this doc; current tree green at `9911b8c`.
|
||||
@@ -1,36 +0,0 @@
|
||||
# Player onboarding data layer §B2: email-OTP, account-link, QR, Bind-Code (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-07-03
|
||||
- **Area:** `internal/api` (onboarding/auth-bind handlers), `internal/store` (migrations 0004–0006)
|
||||
- **Commits:**
|
||||
- `dbe34a1` feat(api): player email-OTP verification (§B2) — `POST /account/email/{start,verify}`; 6-digit code, SHA-256-at-rest, 10-min TTL, 5-attempt cap enforced in the repo
|
||||
- `1f8b9bb` feat(api): record account-link auth source (`mojang|thirdparty`) for the dual-Yggdrasil split (§10)
|
||||
- `116595f` feat(api): QR scan-login completion poll on the internal face (`GET /internal/account/link/status/{mc_uuid}`) — read-only, reuses `UserByMCUUID`, no migration
|
||||
- `fe2ece0` feat(api): public Bind-Code onboarding (`POST /auth/bind`) — the one pre-account entrypoint of `console.<root_domain>`; refuses a staff-UUID code with 403 without consuming it, so the public door provably never yields an admin principal
|
||||
- `55592ed` feat(auth): public auth-bind endpoint wiring
|
||||
- `6c3999a` fix(api): rate-limit email-OTP sends to close the email-bomb vector
|
||||
- `879b177` fix(api): make OTP-start throttle atomic to close the concurrent-burst bypass
|
||||
- **Tasks:** #29 (B2 data layer), #32 (OTP rate-limit), #35 (atomic throttle), #39 (console access model)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the Go-verifiable data layer of forced web onboarding: prove control of an email
|
||||
(OTP), record which Yggdrasil authenticated an in-game UUID, let a phone already signed in
|
||||
to the panel complete a QR device-code link, and let an account-less player redeem a
|
||||
one-time Bind Code minted in the Login Lobby to create+link+session in one public step.
|
||||
The two fixes bound the OTP abuse surface — a per-target send rate limit and an atomic
|
||||
reserve that closes the check-then-act race on the attempt counter.
|
||||
|
||||
## Why
|
||||
|
||||
The spec forces onboarding through the web so every account is provably email-controlled
|
||||
and UUID-linked before it can operate anything. The `op.console` redline in `fe2ece0` — a
|
||||
staff-UUID code is refused without being consumed — is what keeps the public console door
|
||||
from ever minting an admin principal.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The account/session
|
||||
> logic, single-use codes, and the op.console redline were covered by handler tests + the
|
||||
> OpenAPI parity gate at each commit; the identity guarantee behind a Bind Code lives in
|
||||
> velocity/Java (CODE-ONLY) and is not verifiable from this repo. Not independently
|
||||
> re-verified for this doc; current tree green at `9911b8c`.
|
||||
@@ -1,32 +0,0 @@
|
||||
# §B3 username-collision reclaim + account migration (ledger backfill)
|
||||
|
||||
- **Type:** feature — retroactive ledger entry
|
||||
- **Date:** 2026-06-27 – 2026-07-05
|
||||
- **Area:** `internal/api` (internal-face reclaim/blacklist, account migrate), `internal/store` (migration 0006)
|
||||
- **Commits:**
|
||||
- `a29571d` feat(api): reclaim squatted usernames for Mojang-priority players (§B3, 正版优先) — `POST /internal/player/reclaim` bars the squatter UUID + stashes its data (30-day hold) in one transaction, idempotent, returns the *first* reclaim's expiry; `GET /internal/player/blacklist/{mc_uuid}` is the login-gate check
|
||||
- `fdb6efb` feat(account): migrate a live account's owned servers to a new account (§B3 inherit)
|
||||
- **Tasks:** #30 (B3 game-login + username-collision reclaim)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the data layer of the Mojang-priority collision flow: when the configured
|
||||
third-party Yggdrasil and official Mojang issue the same username under different UUIDs,
|
||||
the non-genuine squatter is displaced in favour of the real Mojang owner. Both tables are
|
||||
keyed by `mc_uuid`, so the genuine player — identical username, *different* UUID — is
|
||||
never caught by the bar. `fdb6efb` adds the inherit half: migrating an existing account's
|
||||
owned servers onto a new account.
|
||||
|
||||
## Why
|
||||
|
||||
Two players cannot hold one username across two Yggdrasils; the spec resolves it in the
|
||||
genuine Mojang owner's favour with a 30-day data hold for the displaced squatter, told the
|
||||
truth about how long their data is kept (the first hold's window, never a fresh `now()+30d`
|
||||
on retry).
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Handlers + the
|
||||
> in-memory repo contract were unit-tested at each commit; the Postgres SQL path is
|
||||
> integration-only, and the velocity collision-routing / limbo prompt / authlib
|
||||
> dual-backend are code-only (Java) and out of this data-layer slice. Not independently
|
||||
> re-verified for this doc; current tree green at `9911b8c`. Related: the operator-facing
|
||||
> `/felis migrate` command has its own doc ([felis-migrate-command](2026-07-05-felis-migrate-command.md)).
|
||||
@@ -1,39 +0,0 @@
|
||||
# felis-api security + robustness hardening (audit sweep) (ledger backfill)
|
||||
|
||||
- **Type:** fix — retroactive ledger entry
|
||||
- **Date:** 2026-06-30 – 2026-07-01
|
||||
- **Area:** `internal/api` (login, request-id, listeners, SSE relays, quota/claim, MyServers), `internal/operator`
|
||||
- **Commits:**
|
||||
- `7a51c1d` fix(api): bound concurrent login bcrypt to shed CPU-pin floods (429 `auth_busy` before the compare; a cap, not a per-account lockout) — *audit #2*
|
||||
- `164ac44` fix(api): validate inbound `X-Request-Id` before echo + audit persist (≤64 bytes, log-safe charset) — *audit-integrity*
|
||||
- `c6c0772` fix(api): read/idle timeouts on all three listeners via a `newAPIServer` factory (closes Slowloris via `ReadHeaderTimeout`; `WriteTimeout` left unset so SSE isn't severed) — *audit #3*
|
||||
- `3c1d647` fix(api): per-principal SSE stream cap (429 `too_many_streams`) — *audit #1, blast-radius bound*
|
||||
- `d6e3189` fix(api): per-write deadline on SSE relay to sever a stalled reader (the real leak close behind the cap) — *audit #1*
|
||||
- `8f41a00` fix(api): clear the SSE write deadline on return so it can't leak onto a reused keep-alive connection — *audit #1*
|
||||
- `6368ab1` fix(api): `COALESCE` the MyServers `owned` flag so an ownerless row doesn't 500 the listing
|
||||
- `2a4a81b` fix(api): don't burn the wake cooldown when refused at capacity
|
||||
- `9873904` fix(operator): populate `Status.Players` from an RCON `list` probe (so the panel doesn't report 0/0)
|
||||
- **Tasks:** #33 (wake cooldown), #34 (Status.Players), #41–#46 (audit #1–#4)
|
||||
|
||||
## What it did
|
||||
|
||||
A hardening sweep across the API's abuse and robustness surface: bound the two unbounded
|
||||
CPU/goroutine amplifiers (concurrent bcrypt, per-principal SSE streams), close the SSE
|
||||
relay's real stalled-reader leak with a per-write deadline (and clear it so it can't leak
|
||||
onto a pooled connection), validate the caller-supplied request id before it reaches the
|
||||
audit trail, set listener timeouts to close Slowloris, and fix two functional bugs — the
|
||||
ownerless-row 500 and the wake cooldown burned on a capacity refusal.
|
||||
|
||||
## Why
|
||||
|
||||
Each is a specific, demonstrated failure mode: a login flood pins every core in bcrypt; a
|
||||
stalled SSE reader leaks a relay goroutine + its upstream kube-apiserver follow *for the
|
||||
life of the process*; an unvalidated `X-Request-Id` is a CR/LF log-forgery vector. The
|
||||
`WriteTimeout`-left-unset detail is load-bearing — a blanket write timeout would sever the
|
||||
healthy long-lived console/build-log streams the platform depends on.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Each fix shipped a
|
||||
> targeted test at its commit — notably `d6e3189`/`8f41a00` use a deadline-aware
|
||||
> `ResponseWriter` that fails closed if the guard is removed. The quota-claim TOCTOU
|
||||
> (audit #4) is a documented KNOWN-LIMITATION (`2c56d17`), closeable only against a real
|
||||
> Postgres. Not independently re-verified for this doc; current tree green at `9911b8c`.
|
||||
@@ -1,25 +0,0 @@
|
||||
# felis_* Prometheus metrics (§23) (ledger backfill)
|
||||
|
||||
- **Type:** feature — retroactive ledger entry
|
||||
- **Date:** 2026-06-30
|
||||
- **Area:** `internal/metrics` + the emit sites in build, platform/fleet, and the start lifecycle
|
||||
- **Commits:**
|
||||
- `75642d9` feat(metrics): named `felis_*` Prometheus collectors
|
||||
- `2a93a9e` feat(metrics): record `felis_image_build_failures_total` on failed builds
|
||||
- `79eae7f` feat(metrics): publish `felis_servers_total` from a fleet snapshot
|
||||
- `8ac5e64` feat(metrics): observe `felis_start_duration_seconds` across the start lifecycle
|
||||
- **Tasks:** #17 (§23 felis_* metrics decision)
|
||||
|
||||
## What it did
|
||||
|
||||
Added the named `felis_*` collector set and wired the three emit points that make it
|
||||
non-empty: a counter incremented on image-build failure, a gauge published from a fleet
|
||||
snapshot, and a histogram observed across the server start lifecycle.
|
||||
|
||||
## Why
|
||||
|
||||
§23 calls for first-class operational metrics under a stable `felis_` namespace rather than
|
||||
ad-hoc logging, so an operator can alert on build failures, fleet size, and start latency.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. Not independently
|
||||
> re-verified for this doc; current tree green at `9911b8c` (WSL oracle, go1.26.4).
|
||||
@@ -1,37 +0,0 @@
|
||||
# Auto-update subsystem: decision core + sources + gatherer + window API (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-07-01 – 2026-07-05
|
||||
- **Area:** `internal/updates` (pure decision core), `internal/updater` (release sources, gatherer), `internal/api` (window admin API)
|
||||
- **Commits:**
|
||||
- `c01f133` feat(updates): pure I/O-free decision core — each tracked component is Pinned (Minecraft, left alone), Notify, or Scheduled (apply only inside a SysAdmin window); never force-applied, never a downgrade, never an auto-applied prerelease
|
||||
- `3673af6` feat(api): admin API for the maintenance window (`GET`/`PUT /updates/window`), stored as JSON under `platform_settings` — API + persistence only, nothing consumes it yet
|
||||
- `7464fa7` fix(updates): tag `Window` JSON so the persisted window round-trips (the obvious decode is correct by construction; a zero window fails closed to notify-only)
|
||||
- `96b3cc9` feat(updater): wire `updates.Run` to a caller with PaperMC v3 release discovery
|
||||
- `7d27640` feat(updater): GitHub Releases source, routing felis-api/k3s/cloudflared
|
||||
- `7db57b9` feat(updater): `VersionGatherer` extraction core + CLI gather seam
|
||||
- **Tasks:** #38 (auto-update: Felis/k3s/components/Velocity, pin Minecraft)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the auto-update spine as a pure decision core plus the release-discovery sources
|
||||
(PaperMC, GitHub Releases) and the version gatherer, with a SysAdmin-set maintenance
|
||||
window read/written through an admin API. Version parsing tolerates the real feeds (leading
|
||||
`v`, k3s `+k3s1` suffix, calendar versions, prerelease tails) and orders by SemVer
|
||||
precedence.
|
||||
|
||||
## Why
|
||||
|
||||
The red lines are `不要强制自动更新` (never force auto-update) and `能不动的就别动`
|
||||
(Minecraft stays pinned). The design encodes them structurally: a component may be applied
|
||||
*only* inside a window the operator explicitly set, and Minecraft is Pinned so it is never
|
||||
touched. `7464fa7`'s fail-closed zero-window (decodes to notify-only, never a rogue apply)
|
||||
is the safety property for the not-yet-built runner.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The load-bearing
|
||||
> invariants (pinned never changes, no downgrade, no auto-prerelease, apply-only-in-window)
|
||||
> and the JSON round-trip contract were unit-tested at their commits. This subsystem is
|
||||
> deliberately **report-only / integration-deferred**: the concrete Notifier/Applier,
|
||||
> the `felis update` CLI + CronJob, and the current-version producing seams are declared
|
||||
> but not wired (see `internal/updater/doc.go`, `openapi.yaml`). Not independently
|
||||
> re-verified for this doc; current tree green at `9911b8c`.
|
||||
@@ -1,37 +0,0 @@
|
||||
# Passkey (WebAuthn) enrollment subsystem + hardening (ledger backfill)
|
||||
|
||||
- **Type:** feature + fix — retroactive ledger entry
|
||||
- **Date:** 2026-07-01 – 2026-07-02
|
||||
- **Area:** `internal/passkey` (go-webauthn adapter), `internal/api` (enrollment handlers/audit), `internal/store` (migrations 0007–0009)
|
||||
- **Commits:**
|
||||
- `f2c916d` feat(api): passkey enrollment persistence layer
|
||||
- `742f15f` feat(api): passkey enrollment endpoints
|
||||
- `0261204` feat(passkey): go-webauthn enrollment verifier adapter (Oracle-verified against a virtual authenticator)
|
||||
- `fce0fce` feat(passkey): wire the enrollment verifier into felis-api
|
||||
- `7278cd7` feat(passkey): require + record user verification at enrollment (`UserVerification=required`; capture `user_verified`/`backup_eligible`/`backup_state` — migration 0009) — *fix (d)*
|
||||
- `cdbb5ab` fix(api): record credential id in the passkey-register audit event so bind/unbind are symmetric — *fix (a)*
|
||||
- `9953275` fix(api): bound `webauthn_challenges` growth by superseding *all* prior rows per (user, purpose) — *fix (b)*
|
||||
- `20e31fb` fix(store): cascade-delete passkeys + challenges on user removal (recreate both FKs `ON DELETE CASCADE`, scoped to the passkey tables only) — *fix (c)*
|
||||
- `54bc6ef` fix(api): clear bound passkeys on password change to close a takeover foothold — *fix (e)*
|
||||
- **Tasks:** #36 (passkey bind with email-OTP fallback), #48–#52 (fixes a–e)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the WebAuthn *enrollment* half — persistence, the go-webauthn crypto adapter, and
|
||||
the register-begin/finish endpoints — then hardened it through the five-fix batch (a–e):
|
||||
symmetric audit, a bounded challenge table, cascade cleanup, enforced+recorded user
|
||||
verification, and unbinding every passkey on a password reset so a passkey planted through
|
||||
a transiently-hijacked session cannot survive as a standing login foothold.
|
||||
|
||||
## Why
|
||||
|
||||
Passkeys are the phishing-resistant factor with email-OTP as the fallback. The hardening
|
||||
batch closes the seams that make enrollment safe to *rely on*: without UV enforcement a
|
||||
passkey proves possession but not user; without the password-reset clear, a planted
|
||||
passkey outlives the very remediation meant to evict an attacker.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The adapter crypto
|
||||
> was verified against a virtual authenticator (virtualwebauthn), and each fix shipped
|
||||
> with a targeted test (UV-negative rejection, challenge-growth bound, cascade, symmetric
|
||||
> audit) at its commit. Not independently re-verified for this doc; current tree green at
|
||||
> `9911b8c`. The assertion/login half is a separate doc ([passkey-login](2026-07-01-passkey-login.md)).
|
||||
@@ -1,36 +0,0 @@
|
||||
# Passkey (WebAuthn) login: assertion, discoverable, clone-detection (ledger backfill)
|
||||
|
||||
- **Type:** feature — retroactive ledger entry
|
||||
- **Date:** 2026-07-01 – 2026-07-05
|
||||
- **Area:** `internal/passkey` (assertion crypto), `internal/api` (login/assertion, unbind, UA-guard), `internal/store` (migrations 0013/0014)
|
||||
- **Commits:**
|
||||
- `e035142` feat(passkey): WebAuthn login/assertion crypto adapter (BeginLogin/FinishLogin over go-webauthn, Oracle-verified against a virtual authenticator; surfaces the signature counter as a ceremony fact)
|
||||
- `ec468ba` feat(auth): discoverable (usernameless) passkey login — the from-zero door the username-first assertion couldn't key on
|
||||
- `0dbd557` fix(store): renumber the discoverable-login migration 0013 → 0014
|
||||
- `9e1df12` feat(passkey): advance `sign_count`, reject clone-warned assertions
|
||||
- `4f59d51` feat(auth): owner-tier passkey-unbind remediation endpoint
|
||||
- `a63f49d` feat(panel): steer WeChat/QQ in-app browsers to the system browser for passkey — a backend-only UA interstitial (the SPA is untouched); asset/API/health requests pass through, an `ua_ack` cookie lets a determined user continue
|
||||
- **Tasks:** #40 (from-zero discoverable login), #67 (WeChat/QQ UA-guard in `internal/panel`)
|
||||
|
||||
## What it did
|
||||
|
||||
Built the assertion (login) half of the ceremony: the crypto adapter, then discoverable
|
||||
credentials so a user with no typed identifier can still log in (the enrollment
|
||||
identifier problem the earlier deferral doc named), clone detection via the advancing
|
||||
signature counter, and the owner-tier unbind remediation. `a63f49d` guards the flow at the
|
||||
transport edge — WebAuthn is unusable inside the WeChat/QQ WebViews, so those UAs get a
|
||||
bilingual "open in your system browser" page instead of the passkey SPA.
|
||||
|
||||
## Why
|
||||
|
||||
Enrollment without a login path is half a feature. Discoverable credentials resolve the
|
||||
blocker recorded in the earlier deferral (`users.email` is nullable/non-unique and a
|
||||
player's username is their Minecraft UUID, so username-first assertion had nothing to key
|
||||
on). The UA-guard stops the most common real-world dead end: a passkey prompt that can
|
||||
never succeed inside an in-app browser.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The assertion crypto
|
||||
> was verified against a virtual authenticator (enrollment→assertion chain, origin-mismatch
|
||||
> and unbound-credential rejection); `9e1df12`'s clone policy and the UA-guard pass-through
|
||||
> were unit-tested at their commits. Not independently re-verified for this doc; current
|
||||
> tree green at `9911b8c`.
|
||||
@@ -1,37 +0,0 @@
|
||||
# System servers: login-limbo + lobby (always-on gate) (ledger backfill)
|
||||
|
||||
- **Type:** feature — retroactive ledger entry
|
||||
- **Date:** 2026-07-02
|
||||
- **Area:** `internal/config`, `internal/naming`, `internal/api` (CRD readiness), `internal/operator`, `internal/platform`, `cmd/felis`, `plugins/limbo`, `deploy/limbo` + `deploy/lobby`
|
||||
- **Commits:**
|
||||
- `9bed51b` feat(config): `[velocity] login_image/lobby_image` — setup provisions the always-on system services only when set (empty = fail-loud skip; no official LOOHP/Limbo image exists)
|
||||
- `9ef817f` feat(naming): reserved system-server names + service-token identifiers (single source of truth for the internal-API credential Secret)
|
||||
- `159107b` feat(api): HTTP readiness knob on `MinecraftServer` + user-server fallback defaults to the login gate
|
||||
- `dc23cb5` feat(operator): system-server pod HTTP readiness probe + login-only `FELIS_SERVICE_TOKEN` env (keyed off the reserved name so it can never leak into a user pod; sourced via `secretKeyRef`, never inlined)
|
||||
- `3fdb3d0` feat(platform): internal-API base-URL helper + single-sourced token Secret
|
||||
- `f554d52` feat(cli): provision the reaper-exempt login/lobby servers + replicate the service-token Secret into the minecraft namespace
|
||||
- `241fe21` feat(limbo): felis-limbo in-game login flow (join → blacklist check → mint bind code → open book to `console.<root_domain>` → poll link-status → BungeeCord transfer to lobby; fail-closed)
|
||||
- `c7315e4` feat(deploy): login-limbo + lobby images with game-port pinning (server-port pinned to GamePort 25565 on every start)
|
||||
- **Tasks:** #53–#68 (system-server plumbing L1–L4, limbo plugin, operator env injection)
|
||||
|
||||
## What it did
|
||||
|
||||
Stood up the always-on authentication gate: reserved, reaper-exempt login/lobby
|
||||
`MinecraftServer`s provisioned by setup, an HTTP readiness path for the RCON-less LOOHP/Limbo
|
||||
loader (which reports "started" only after the first tick), and the felis-limbo plugin that
|
||||
runs the whole onboarding *inside* Limbo before transferring an admitted player to the
|
||||
lobby. A fresh connection always lands on the login gate, never a user backend, so
|
||||
authentication is always in front.
|
||||
|
||||
## Why
|
||||
|
||||
The spec requires that a player authenticate before reaching any real server. That needs a
|
||||
purpose-built always-on front server (Limbo) that speaks to the internal API — hence the
|
||||
login-only service-token injection (keyed to the reserved name so it can never reach a user
|
||||
pod) and the HTTP readiness knob for a loader that has no RCON.
|
||||
|
||||
> **Backfill note.** Reconstructed 2026-07-07 from the commit history. The Go layer
|
||||
> (config/naming/readiness/operator env/platform) was unit-tested at each commit; the
|
||||
> felis-limbo plugin is Java verified against a real Limbo jar via podman (#65), and the
|
||||
> images carry a real build+boot check (#63, limbo `/healthz` 200 on 25565). Not
|
||||
> independently re-verified for this doc; current tree green at `9911b8c`.
|
||||
Loaded 100 of 390 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user