From 403a751c96348c35d3395b327b403e652c39b67f Mon Sep 17 00:00:00 2001 From: Lemon-miaow Date: Sun, 27 Sep 2026 10:59:20 +0800 Subject: [PATCH] =?UTF-8?q?docs(dr):=20=E9=87=8D=E5=BB=BA=20runbook=20?= =?UTF-8?q?=E8=A1=A5=E9=82=AE=E4=BB=B6=E8=87=AA=E6=A3=80=E3=80=81domain=20?= =?UTF-8?q?set=E3=80=81cloudflared=20=E5=87=AD=E6=8D=AE=E6=81=A2=E5=A4=8D?= =?UTF-8?q?=E4=B8=8E=E9=82=AE=E4=BB=B6=E9=AA=8C=E6=94=B6?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/operations.md | 5 ++- docs/troubleshooting.md | 99 ++++++++++++++++++++++++++++++++++++----- 2 files changed, 91 insertions(+), 13 deletions(-) diff --git a/docs/operations.md b/docs/operations.md index f0278ce..bb3cdf4 100644 --- a/docs/operations.md +++ b/docs/operations.md @@ -598,8 +598,9 @@ production install: holds only sealed objects. - **Keep one database bundle off the host** as well when there is no bucket. It contains `secrets.env`, which a rebuild needs to read the rest. -- **Rehearse the rebuild** once on a spare VM: §16 "Rebuild on a new host", steps 1–5, - then log in and restore one world. `felis offsite status` and `felis db check` exit +- **Rehearse the rebuild** once on a spare VM: troubleshooting.md §16 "Rebuild on a new + host", every step but 8 (take-over) and 11 (the tunnel), then its checks: sign in with + an email code, restore one world and join it. `felis offsite status` and `felis db check` exit non-zero when the copy or the newest bundle is stale; wire them into your monitoring, or rely on the watchdog's mail. diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 5775f2e..5b85771 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -2289,6 +2289,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds: | Live worlds | `world-*` volumes under `/var/lib/rancher/k3s/storage` | **only as their archives** | a restore from the newest archive (§10) | everything since that world's newest archive | | User images | `registry` volume | hourly; image lists kept 14 days | `fetch-images` | images pushed in the last hour | | Submission uploads (modpacks awaiting or past review) | `felis-uploads` volume | hourly; upload lists kept 14 days | `fetch-uploads` | uploads of the last hour | +| The Cloudflare Tunnel connector (when the panel is behind one) | `/etc/felis/cloudflared.yml`; `/root/.cloudflared/cert.pem` and `.json`; `cloudflared-felis.service` | the config inside every bundle; the login certificate, the tunnel's credentials and the unit are not copied | the Cloudflare step of `sudo felis setup` (step 11 below) | nothing: the tunnel, its DNS records and the Access application live at Cloudflare | | Platform images, Velocity, the JRE, build tools | registry, `/opt/felis` | not copied | the installer builds and pushes them again | nothing | | k3s itself (its token, CA, datastore, Secrets, Deployments) | `/var/lib/rancher/k3s` | not copied | the installer makes a new single-node cluster and renders every Secret and Deployment from `/etc/felis` | nothing: no Felis data lives only there | @@ -2313,17 +2314,21 @@ and each skips what is already in place, so an interrupted one resumes. Write those sizes down with the bucket's download rate and you have the recovery time for your install. A whole-host rehearsal on a spare machine, once per release, is the way to know it for sure. The spare follows the steps below -without step 8: it finds the production host named in the bucket and stands -by, so it copies nothing into the bucket, prunes nothing there and, while -production keeps writing, mails none of the owners its restored database -holds. +without steps 8 and 11. It finds the production host named in the bucket and +stands by, so it copies nothing into the bucket, prunes nothing there and, +while production keeps writing, mails none of the owners its restored +database holds. A second connector on the production tunnel would take a +share of the real visitors, so the spare leaves the tunnel alone and is +reached by its own address (step 10 moves it to one). The order below matters: the state goes in before the installer so it reuses the old secrets and bucket; the images go back before the database so the servers the database restores find the digests they pin, inside the pruner's 24-hour grace; the MinecraftServers go back with the database that names their -owners; the worlds come last because a restore needs a server to restore -into. +owners; the worlds come after them because a restore needs a server to +restore into. Mail is checked before the domain moves, because a move can +cost passkeys and email codes are the way back in; the domain moves after the +MinecraftServers are applied, since the login gate's env carries it. ### Rebuild on a new host (the old one is gone) @@ -2441,6 +2446,67 @@ host yourself, plus the off-site encryption key if the copy is in the bucket. wrote it, and exits 4. With `-yes` it records this host; the old host, if it ever runs again, copies nothing more and says it was taken over. Skip this step on a rehearsal machine. +9. Check that this host can send mail, when the old one had a relay. Email + codes are how users sign in without a passkey, and step 10 can cost them + their passkeys. Run `sudo felis setup`, press `e` (configure email) on the + status screen and save the pre-filled relay with its password typed again + (the form never shows the stored one). Saving delivers one self-test + message to the From address and writes nothing unless it arrives. A relay + that admits senders by address, or an SPF record for the From domain that + lists the old host's address, refuses this host or sends its mail to spam: + add the new address there and save again. With no relay at all, an Owner + who cannot sign in recovers with `sudo felis breakGlass` (§17). +10. Move the install to this host's address when its names still lead to the + old one. The installer kept the bundle's root domain (its log says + `(reusing the installed domain)`). + + - The `.nip.io` default resolves to the dead host by + construction. Move it: + + ``` + sudo felis domain set .nip.io # the plan + sudo felis domain set -yes .nip.io + sudo felis domain check + ``` + + It runs after step 6 because the MinecraftServers applied there carry + the domain in the login gate's env. The plan counts the passkeys that + stop working; their users sign in with an email code (step 9) and + register a new one. docs/operations.md §6 has the rest of what it moves. + - A domain of your own stays. Point its records at the new address: + ``, `*.`, `console.` and `op.console.`, or only + `` and `*.` when the tunnel serves the panel. Confirm each + with `dig +short ` before going on: `felis domain check` warns + while a name does not resolve and accepts any address once it does. + Lower the records' TTL beforehand if the old host is still around to + plan with. +11. Bring the Cloudflare edge back, when the old host served the panel + through a tunnel. The bundle carries `/etc/felis/cloudflared.yml`; the + login certificate (`/root/.cloudflared/cert.pem`), the tunnel's + credentials (`/root/.cloudflared/.json`) and the + `cloudflared-felis` unit stay behind with the old disk, so the panel + hostnames answer Cloudflare error 1033 until a connector runs here. + + Run `sudo felis setup`, press `c` (change connection) on the status + screen and choose Cloudflare Tunnel + Access. On the step's first screen + press `i` to install cloudflared, then `l` for `cloudflared tunnel login` + (browser consent on your account, which writes `cert.pem`); `enter` opens + the form once both are in place. Enter an API token and the account ID, + then the same Admit identity, hostnames and tunnel name as before. The + hostnames come pre-filled from the restored config and the tunnel name + defaults to `felis`; for another name, look up the id on the `tunnel:` + line of the restored `cloudflared.yml` in Zero Trust → Networks → Tunnels. + + The step finds the existing tunnel by name and fetches its credentials + again, points the panel records at it, finds the existing Access + application (so its audience stays the one the restored config names) + and rewrites its policy, installs and starts `cloudflared-felis`, and + closes the panel's NodePort once the tunnel serves. A kept copy of the + old `.json` can go back into `/root/.cloudflared` (mode 0600) + first; the step still needs `cert.pem`. A different tunnel name makes a + second tunnel, moves the panel records to it and leaves the old one idle: + delete that one in the dashboard afterwards. Skip this step on a + rehearsal machine. Check the rebuild before letting players in: @@ -2449,13 +2515,24 @@ sudo felis db check # the database answers and has a fresh b kubectl get minecraftservers -A # every server the bundle held kubectl -n minecraft get pods # servers pull their pinned images (no ImagePullBackOff) sudo felis offsite status # the hourly copy runs from this host again +sudo felis domain check # every name, the certificate and the tunnel on the new address ``` -Point the panel and game hostnames at the new host (DNS, or the tunnel in -front of it). In the panel: sign in with an old account (accounts and passkeys -come back with the database), open a restored server's backup page and confirm -its archives are listed, restore the newest one, start the server and join -it. +A rehearsal machine that restored a tunnel install and moved to a `nip.io` +name in step 10 fails the Cloudflare tunnel line, since the tunnel still +routes production's names; that one failure is expected there. + +In the panel, reached by its hostname (steps 10 and 11): + +- Sign in with an old account; accounts and passkeys come back with the + database, except passkeys step 10 counted as lost. +- Sign out and sign in again with an email code, to an Owner account whose + inbox you read. The code arriving proves felis-api itself mails through the + relay to an outside inbox; step 9's self-test came from the setup console + and went to the From address. A code that never comes: §17. +- Open a restored server's backup page and confirm its archives are listed, + restore the newest one, start the server and join it at + `.`. ### Keep a copy somewhere else