Commit Graph
34 Commits
Author SHA1 Message Date
Lemon-miaow 47890ca913 feat(offsite): 世界归档与数据库备份加密同步到异地 S3,reaper 确认异地副本后才删除世界 2026-09-24 19:25:16 +08:00
Lemon-miaow d17524cd67 feat(watchdog): 主机侧巡检定时器按异常邮件通知平台所有者,operator 增加 phase 与 build_info 指标、卡死存活探针与告警规则 2026-09-24 18:06:19 +08:00
Lemon-miaow 215bfd78d7 fix(images): 服务器镜像在创建时固定到仓库 digest,更换镜像需确认备份,安装器重建前先固定旧服并推送不可变版本标签 2026-09-24 16:57:38 +08:00
Lemon-miaow 5521e498a9 fix(platform): minecraft 命名空间强制 PodSecurity baseline,reaper 世界根目录改走静态 hostPath PV 2026-09-24 16:28:12 +08:00
Lemon-miaow c7db7d4126 feat(db): 控制面 PG 定时备份、迁移前快照与原子恢复 2026-09-24 15:19:42 +08:00
Lemon-miaow 7819e5de50 feat(netpol): 锁定游戏服出站并为 registry 加入站围栏 2026-09-24 14:25:17 +08:00
Lemon-miaow 2961beb8fe test(bootstrap): keep the worlds-root warning case hermetic on hosts that already run k3s (#73) 2026-09-24 03:11:10 +08:00
Lemon-miaow b14bacbfc6 fix(build): mirror the Trivy Java DB — jar-bearing builds failed closed at the scan gate (#72) 2026-09-24 03:11:01 +08:00
Lemon-miaow 328e570309 fix(setup,install): refresh the workload namespace's felis-config mirror (#51) 2026-09-23 20:07:20 +08:00
Lemon-miaow 4d3c85fd06 fix(bootstrap): [smtp] carry stops hoarding the auth_source comment block (#50) 2026-09-23 19:45:26 +08:00
Lemon-miaow c7e585e21d fix(bootstrap): mirror the image batch under ONE docker start/stop
Live re-run: the per-image systemctl start/stop docker cycles tripped systemd's
start rate limit after three fast pushes — "Start request repeated too
quickly / start-limit-hit" — and the fourth image (the paper base) silently
never reached the registry while the installer aborted. docker.service is
socket-triggered, so every cycle counts against the burst limit twice.

push_images_to_registry now starts docker once for the whole batch and stops it
once at the end; push_image_to_registry itself no longer touches systemd.
bootstrap_test.sh pins the wrap (exactly one start, one stop, four pushes).
2026-09-23 19:23:44 +08:00
Lemon-miaow b8e554dac7 fix(bootstrap): carry the operator's [archive] keys across re-runs too
Same class as 765a892, same table-level amnesia: [archive] retention /
warn_before / max_local_bytes are the reaper's runtime knobs (read from the
config Secret at job time; built-ins 90d / 3d,1d / no cap), and write_felis_toml
rewrote the whole table as store+local_path on every re-run. An operator who
narrowed the retention window silently got the 90d built-in back.

persisted_archive_block carries the three keys forward; store and local_path
stay installer-owned (FELIS_ARCHIVE_LOCAL_PATH must equal the mount the render
passes). Extends the bootstrap_test carry case with the archive keys and the
installer-owned exclusion.
2026-09-23 19:09:57 +08:00
Lemon-miaow 765a8923a4 fix(bootstrap): re-runs keep the operator's [registry] overrides
§15's upgrade path is "re-run the installer", but write_felis_toml rewrote the
[registry] table from scratch — url + build_namespace only. Everything else an
operator put there (the §8e build-lane executor mirrors, the resource caps, the
uploads backend stamped by the storage wizard, [registry.s3]) was silently
reverted on every re-run: builds went back to the denied upstream executors and
an S3-backed install flipped to local storage, with nothing pointing at why.

Found while landing the registry-hosting work, which depends on those same
keys surviving.

- persisted_registry_block carries the operator-owned [registry] keys and the
  [registry.s3] subtable forward, same first-readable-file rule as
  persisted_smtp_block; url/build_namespace stay installer-owned (they must
  match REGISTRY_URL/BUILD_NS, so a stale value must NOT survive).
- The s3 subtable header is re-emitted with its keys, so nothing carried lands
  as an unknown key under [registry].
- bootstrap_test.sh pins the carry, the installer-owned exclusion, and
  idempotence (a second re-run writes a byte-identical file).
2026-09-23 19:05:54 +08:00
Lemon-miaow 13d64e0000 feat(bootstrap): host every built image in the internal registry — GC-durable pulls
The disk-pressure drill's dead end: kubelet's image GC collects an unused image
and an air-gapped node has nothing to pull it from (ImagePullBackOff until an
operator re-imports). The registry the bundle already renders becomes that pull
source:

- Every image the installer builds is now a registry ref
  (registry.felis.svc:5000/felis/{felis,limbo,lobby,paper}:demo), imported into
  containerd under that exact name (first boot needs no registry round-trip)
  and mirrored into the registry after deploy_bundle (push_image_to_registry:
  push endpoint 127.0.0.1:5000, and only the path after the host matters to the
  registry — a push there lands where kubelet's mirrored pull looks). A ref
  outside the registry is warned about, not silently unmirrored.

- configure_registry_mirror writes /etc/rancher/k3s/registries.yaml mapping
  registry.felis.svc:5000 onto http://127.0.0.1:5000, the loopback hostPort the
  registry Deployment binds (node containerd cannot dial the Service VIP — live
  drill: "Empty reply"). k3s regenerates containerd config only at agent start,
  so a CONTENT change restarts k3s and an identical file (every re-run)
  restarts nothing.

- import_registry_image caches registry:2 into containerd so the registry
  Deployment can start on a box that cannot reach Docker Hub.

- Migration 0021 re-points the recommended whitelist seeds ('felis-lobby:demo',
  'felis-paper:demo') at the registry refs — a user server created from those
  rows must not strand when GC collects the bare tag. Only recommended rows
  still holding the old seed are touched; enabled is preserved; a pre-existing
  target row wins over a duplicate.

bootstrap_test.sh pins the mirror idempotence (identical content must NOT
restart k3s), the push-ref mapping (including the port-confusion refusal) and
the registry:2 precheck.
2026-09-23 19:02:58 +08:00
Lemon-miaow 2b87a5a13b fix(install): grant the reaper traverse on the worlds root (#6)
Live drill found this: the reaper pod runs as uid 1000, k3s creates its
storage root /var/lib/rancher/k3s/storage 0700 root:root, so enabling
retention on a stock install made every archive fail
'lstat /worlds/<pvc>: permission denied' and skip the world (fail-closed,
but a silent no-op). bootstrap now grants traverse (setfacl u:1000:x,
else chmod o+x) when FELIS_WORLDS_HOST_PATH is set, the renderer's
precondition note names the requirement, and troubleshooting documents
both it and the multi-node nodeSelector fact.

Verified on the VM after granting the ACL: a 20d-idle world with a marker
file was archived into felis-backups (marker intact), its PVC and host
directory were reclaimed, world_backups got an inactive_15d row, and the
servers row/CR were retained.
2026-09-22 21:03:17 +08:00
Lemon-miaow fd33fd05e1 fix(install): backups exist on a default install; retention resolves real world dirs (#6)
Three faces of one gap, all on the supported install path:

- Backup/restore answered 503 out of the box: nothing ever rendered the
  archive PVC, so FELIS_BACKUP_PVC was unset. The bundle now renders the
  PVC (Minecraft namespace, RWO 10Gi, cluster default class) and
  'felis manifests' names it by default (--backup-pvc= is the explicit
  no-store shape); bootstrap passes it through so the generated felis.toml
  [archive] local_path and the jobs' mount path come from one variable.
- Retention was unreachable: bootstrap never passed the reaper flags. It
  now forwards FELIS_WORLDS_HOST_PATH/FELIS_ARCHIVE_LOCAL_PATH, so one
  env enables the daily CronJob; unset keeps today's fail-safe (no reaper,
  nothing deleted).
- Even when enabled it could not find a world on a stock install:
  resolveWorldDir now also resolves the exact local-path directory
  <pv-name>_<ns>_<pvc-name> read from the live PVC's volumeName (never a
  glob, so a stale deleted PV's bytes can't be archived in place of the
  current world). Reaper Role gains persistentvolumeclaims:get (weaker
  than the delete it already held).

README (zh/en) stops promising automatic/scheduled backups and states
retention is opt-in. bootstrap_test covers the env->flag contract.
2026-09-22 20:48:07 +08:00
flyemoji 7b5b28c587 fix(bootstrap): keep the nano build toolchain under /opt/felis
The source build of the nano binary installed Go at /usr/local/go and
replaced whatever version was already there. On a host that also
builds other things, the operator's own toolchain was removed and
swapped for Felis's pinned version without a word.

GOROOT_DIR is now /opt/felis/go, next to the source, the Velocity
install and the JRE Felis already keeps under /opt/felis, and
install_go_toolchain creates the parent before unpacking. A host where
an earlier run put Go at /usr/local/go downloads it once more on the
next re-run and keeps the old tree untouched; removing it is the
operator's call. The harness now requires the toolchain directory to
be under /opt/felis.
2026-09-22 14:59:39 +09:00
flyemoji 0faec2b02a fix(bootstrap): open the nano port to the proxy alone
For a non-loopback bind, configure_nano_firewall opened the nano port
in firewalld to every source, while the summary told the operator to
restrict it to the proxy. hasJoined takes no token, so on a public
host that port is an auth relay anyone can point a proxy at, spending
this host's Mojang egress until Mojang rate-limits it and the
operator's own players stop getting in.

A new FELIS_NANO_PROXY_CIDR names the proxy. With it, firewalld gets
one rich rule that admits the port from that source only, ipv4 or
ipv6 by the address given. Without it, no port is opened and the
summary prints the rule to add. A re-run closes the port an earlier
installer opened to every source. A rule for a previous
FELIS_NANO_PROXY_CIDR is not tracked and stays until removed by hand.
Hosts without firewalld are handled as before.

The value goes into the rule text, so it is checked up front for an
address with one prefix length and nothing else. firewalld's own
parser accepts both rule forms and refuses an ipv6 address under the
ipv4 family. The harness covers the rule for each family, the
closed-by-default case, the re-run cleanup, the loopback case and the
CIDR check.
2026-09-22 14:51:49 +09:00
flyemoji fa7b54f5ab fix(bootstrap): refuse an unbracketed ipv6 nano listen address
validate_listen checked only the port, so FELIS_NANO_LISTEN=::1:8081
passed. Go refuses that form ("too many colons in address") and needs
[::1]:8081, so the unit crash-looped on every start. A host part that
contains a colon must now be in brackets.

With that, the bare ::1 pattern in nano_listen_is_loopback can no
longer match an address that gets this far, so it goes. [::1] stays.
The harness adds ::1:8081 to the refused addresses, and [::]:8081 and
:8081, both of which Go binds, to the accepted ones.
2026-09-22 14:08:11 +09:00
flyemoji b58c20311c test(bootstrap): pin the nano listen default to loopback
nano_listen_is_loopback decides whether configure_nano_firewall opens
the port, and hasJoined takes no token. A default that does not
classify as loopback would make every fresh nano host a public auth
relay.

The harness now runs the classifier on four loopback binds and three
routable ones, and feeds it the default resolve_nano_listen applies on
a first install, with no operator value and no existing unit. Setting
that default to 0.0.0.0:8081 or :8081, or counting 0.0.0.0 as
loopback, now fails the harness. Test only.
2026-09-22 13:54:31 +09:00
flyemoji cf65ffdae5 fix(bootstrap): open up a nano-only config dir an older run left 0750
write_nano_config creates a missing /etc/felis as 0755, but it left
an existing one alone. On a nano-only host an older installer made
that directory with a bare mkdir -p, so under a root umask of 027 it
is 0750. The DynamicUser unit cannot search it, so felis-nano cannot
read its config, and a re-run stops at the service check instead of
repairing the directory.

An existing directory is now set to 0755 unless it holds the full
install's secrets.env or bootstrap.done. The full install locks the
directory to 0700 and writes secrets.env right after, so its directory
keeps that mode, and install_nano_service still reports the lockout
rather than this widening it. The mode cases run only where chmod
works; on a filesystem that ignores it the harness skips them.
2026-09-22 13:54:18 +09:00
flyemoji 6794e66c4d fix(bootstrap): fail a tokenless private clone instead of prompting
A source build against a private repository with no FELIS_GITHUB_TOKEN,
or a wrong one, made git ask for a username on /dev/tty, and a piped
install sat there waiting.

git_auth now runs git with GIT_TERMINAL_PROMPT=0 on both arms, so git
fails at once with "terminal prompts disabled". Both fetch_source
failures name FELIS_GITHUB_TOKEN in their message: the fresh clone,
and the fetch into an existing checkout, which had no message of its
own before.
2026-09-22 13:53:03 +09:00
flyemoji 34f73ba19f fix(bootstrap): detect a missing terminal by opening /dev/tty
prompt_install_mode guarded its prompt with `[ ! -r /dev/tty ]`, which
never fires on Linux: /dev/tty is mode 0666 whether or not the process
has a controlling terminal, and only opening it fails. Without a
terminal the menu was printed, the read failed with "No such device or
address", and the default was taken by accident rather than by the
documented path.

The guard now opens /dev/tty in a subshell and takes the "no terminal
for a prompt" path when that fails.
2026-09-22 13:52:51 +09:00
flyemoji 02c079c893 fix(bootstrap): verify the go toolchain tarball against a pinned digest
install_go_toolchain downloaded the tarball to a fixed /tmp name and
unpacked it into /usr/local as root, with no digest check. Another
local user could plant that file first, and nothing would notice a
tampered download.

The tarball is now staged in a mktemp -d directory that the exit
cleanup removes, and its sha256 must match before the old toolchain is
touched, so a refusal leaves the host as it was. The default 1.26.4
carries pinned amd64 and arm64 digests next to its version; they are
the ones https://go.dev/dl/?mode=json&include=all publishes. Any other
FELIS_GO_VERSION has to bring its own FELIS_GO_SHA256, documented in
the header, because no pin can cover a version chosen at run time.
Where and which version gets installed is unchanged.
2026-09-22 13:52:29 +09:00
flyemoji 3b0fc7a3e0 fix(bootstrap): print the address nano binds in the install summary
summary_nano printed the node's primary IP for every non-loopback bind
and 127.0.0.1 for every loopback one. A bind to a second private
address, or to [::1], handed the operator a hasJoined URL that nothing
listens on.

The host is now the part of FELIS_NANO_LISTEN before the last ':'. The
node's IP is used only for a wildcard bind (empty, 0.0.0.0 or [::]),
which names no address a proxy could dial. The loopback and
public-bind notes are unchanged.
2026-09-22 13:52:18 +09:00
flyemoji 404d1172a6 fix(bootstrap): refuse a nano listen address without a usable port
FELIS_NANO_LISTEN was never checked. A bare 8081 opened port 8081 in
the firewall while nano bound nothing, a bare 127.0.0.1 printed
http://127.0.0.1:127.0.0.1/... in the summary, and the unit
crash-looped either way.

validate_settings now requires a ':' and a decimal port of 1-65535
after the last one. It runs after resolve_nano_listen, so an address
read back from an existing unit is checked too, and the default always
passes. [::1]:8081 and 0.0.0.0:8081 are accepted.
2026-09-22 13:52:08 +09:00
flyemoji 3918a4b11a fix(bootstrap): install only the full control plane under felis setup
prompt_install_mode also runs inside felis setup. Setup then goes on
to the Owner and edge setup, which need the control plane, so choosing
nano there always ended in a setup error.

Under felis setup the mode is now full before any prompt or default is
considered, and an explicit FELIS_INSTALL_MODE=nano stops with a
message pointing at deploy/bootstrap.sh. That leaves the felis setup
branch of acquire_nano_binary unreachable, so it goes.
install_embedded_binary stays, since the full install still uses it.
2026-09-22 13:52:00 +09:00
flyemoji 515c4a6496 fix(bootstrap): keep a nano host's listen address and mode on re-run
Re-running the installer is how a nano host updates. That re-run reset
FELIS_NANO_LISTEN to 127.0.0.1:8081, so a proxy on another machine lost
its endpoint and every login through it failed. It also offered the
full control plane as the default, which on a nano host means k3s and
Postgres nobody asked for.

The listen address is now settled by resolve_nano_listen, the first
step of main, so the later checks see the result. The operator's value
wins, then the -listen argument of the installed felis-nano unit, then
loopback. The install mode defaults to nano, at the prompt and without
a terminal, when the felis-nano unit exists and the full install's
bootstrap.done marker does not. Only the full install writes that
marker.

The harness reads back the unit it wrote earlier, and checks the mode
default on a nano-only host, a host with the full install, and a fresh
host.
2026-09-22 13:51:52 +09:00
flyemoji 17b4396460 fix(bootstrap): carry auth_source tables with spaced or quoted headers
A re-run copies the operator's [[auth_source]] tables from the existing
felis toml into the new one. The awk program that finds them matched
only the literal header [[auth_source]], so a table written as
[[ auth_source ]], [["auth_source"]] or [['auth_source']], all valid
TOML, was taken for some other section and dropped from the config.

Each section header now decides afresh whether it opens an auth_source
table, through one regex that allows inner whitespace and a single- or
double-quoted key. The single quote is spelled \047, which gawk and
mawk both honour inside a bracket expression. The harness carries each
spelling and checks that the table still stops at the next section.
2026-09-22 13:51:45 +09:00
flyemoji a0f54df2a6 fix(bootstrap): fail the nano install when the unit does not stay up
install_nano_service printed "enabled and started" straight after
systemctl restart, which returns as soon as the process is forked. An
upgrade that keeps an old felis.toml the new binary rejects (an
[[auth_source]] without a prefix, say) left the unit crash-looping in
auto-restart while the installer reported success, and every login
through the proxy failed.

The install now waits two seconds and asks systemctl is-active. A unit
that exited is in "activating (auto-restart)", which is-active does not
count as active; on real systemd a unit whose process exits 1 under
Restart=on-failure reads activating/auto-restart and is-active returns
non-zero, while a running one reads active/running and returns 0. On
failure the install prints the unit's last 20 journal lines and stops.
This also surfaces a nano unit locked out of an existing 0700 /etc/felis.

The harness runs the extracted function with systemctl stubbed both
ways. Without the check, the dead-unit cases fail.
2026-09-22 13:05:42 +09:00
flyemoji 26f685be0e fix(bootstrap): create the nano config dir world-searchable
felis-nano runs as a systemd DynamicUser, so it can read
/etc/felis/felis.toml only if others may search /etc/felis.
write_nano_config made the directory with a bare mkdir -p, which takes
its mode from root's umask. On a host hardened to umask 027 that is
0750: nano exits on "permission denied", the unit restarts every five
seconds, and no login gets through.

A missing directory is now created 0755 explicitly. An existing one
keeps its mode, because the full install sets it to 0700 to protect its
secrets and widening that from the nano path would expose them. A nano
unit locked out that way is left for the install to report.

The harness runs the extracted function under umask 027 and checks both
cases. Reverting to the bare mkdir fails the first; an unconditional
chmod 0755 fails the second. The mode checks skip on filesystems that
ignore chmod, such as Git Bash on NTFS.
2026-09-22 13:04:34 +09:00
flyemoji 07bafebf0d fix(bootstrap): keep the operator's auth sources across re-runs
write_felis_toml regenerates felis.host.toml and felis.pod.toml with a
wholesale `cat >`, and the [[auth_source]] list was a literal LittleSkin
block in that heredoc. Re-running the installer, which is also what
`felis setup` does, threw away any edit to the list: a root the operator
added stopped admitting logins, and a root they removed came back. The
generated comment invited exactly that edit.

Carry the tables forward the way [smtp] already is: read every
[[auth_source]] table from the existing felis.host.toml (falling back to
felis.pod.toml) and emit the LittleSkin default only when there is no
earlier file at all. An earlier file with no tables stays empty, because
that is a Mojang-only server rather than a missing value; felis-api now
treats an empty list that way.

The file header and the comment above the list now say what survives a
re-run, and point at felis.host.toml, which is what the next run reads.

bootstrap_test.sh extracts the new function from bootstrap.sh and checks
the fresh-install default, an operator's own table carried without the
default or the following section, an empty list staying empty, and the
indented form the setup TUI writes. It passes under dash with gawk and
with mawk; forcing the function to always return the default fails five
of the new cases.
2026-09-22 12:49:25 +09:00
flyemoji d9246ddae6 feat(bootstrap): verify Paper and Velocity jars against Fill's digest 2026-08-04 17:36:47 +09:00
flyemoji 584d31fc49 feat(bootstrap): refuse to install an unverified Velocity fork jar
FELIS_VELOCITY_FORK_JAR replaces the proxy every player connects through, and the
only thing checked about it was that the path pointed at a readable file. A
truncated copy, a stale build left at the same path, or the two-patch jar where
the three-patch one was meant all installed silently.

It now requires FELIS_VELOCITY_FORK_JAR_SHA256 and refuses on a mismatch, hashing
stdin rather than the path for the reason install_via_plugins already documents:
sha256sum escapes its output line for a filename carrying a backslash or newline,
and the leading "\" that adds fails every comparison. The absent-digest refusal
prints the jar's actual hash, so the first run after a deliberate rebuild is one
copy-paste rather than an investigation.

The comparison ignores case and internal spaces. The fork is built on a developer
machine, which is usually Windows, and nothing there prints a digest the way
sha256sum does: Get-FileHash returns uppercase and certutil has shipped the bytes
space-separated. Comparing raw would refuse two of the three spellings of the
correct answer and word the refusal as tampering.

No digest is hardcoded, which is the half of the request this does not deliver.
The fork is built from Felis-Legacy and has never been reproduced on a second
machine, so a constant here would pin one machine's output rather than the fork.
The comment that previously asserted the build "is not byte-reproducible" is gone
too -- it was stated more confidently than the evidence supports. The fork jars on
disk carry Gradle's constant 1980-02-01 entry timestamps, so the usual reason a
jar differs between builds is already absent; that is not proof it reproduces, and
neither claim should sit in the script unmeasured.

This is deliberately not a supply-chain signature and the comment says so: an
operator who can write the jar can write the digest. What it buys is that a path
stops being an identity, and that every later re-run re-checks the same build.

Scope: the fork jar only. The else branch still curls stock Velocity from PaperMC
with no verification at all, and that is the branch a default install takes. The
digest is already in hand there -- Fill v3 returns checksums.sha256 and its
download URL is content-addressed on that same value -- and papermc_latest_jar
discards it. Left alone rather than widened into this change.

deploy/bootstrap_test.sh covers the gate's two refusals, its happy path, and the
two Windows digest spellings. Each case extracts the block under test out of
bootstrap.sh with awk and runs it with die/log stubbed, rather than transcribing
it -- a transcribed copy passes forever after someone edits the original. The
extraction is length-bounded: awk runs an unmatched end pattern to EOF, which
would quietly feed the rest of bootstrap.sh to the shell under test. bootstrap.sh
itself cannot run here; it wants root, a package manager and k3s.

A `shell` CI job runs that plus a syntax check over every tracked script. The
syntax step dispatches on each file's shebang instead of running `sh -n` across
the board. The blanket form looks fine and is a false green: on a developer
machine `sh` is usually bash and accepts everything, while the runner's `sh` is
dash. Verified against the real thing rather than an approximation -- inside
ubuntu:24.04, where /bin/sh is /usr/bin/dash, the dispatching loop passes all six
scripts and the blanket loop dies at bootstrap.sh:191 on the first of its 14
arrays.

Refs: Felis-Legacy #19
2026-07-28 18:01:39 +09:00