docs(dr): 重建 runbook 补邮件自检、domain set、cloudflared 凭据恢复与邮件验收

This commit is contained in:
Lemon-miaow committed 2026-09-27 10:59:20 +08:00
1 parent 2d04e0e637
commit 403a751c96
2 files changed
+91 -13

No files matched your search

+88 -11
View File
@@ -2289,6 +2289,7 @@ off-site bucket (next sections) plus a fresh install. What the host holds:
| Live worlds | `world-*` volumes under `/var/lib/rancher/k3s/storage` | **only as their archives** | a restore from the newest archive (§10) | everything since that world's newest archive |
| User images | `registry` volume | hourly; image lists kept 14 days | `fetch-images` | images pushed in the last hour |
| Submission uploads (modpacks awaiting or past review) | `felis-uploads` volume | hourly; upload lists kept 14 days | `fetch-uploads` | uploads of the last hour |
| The Cloudflare Tunnel connector (when the panel is behind one) | `/etc/felis/cloudflared.yml`; `/root/.cloudflared/cert.pem` and `<tunnel-id>.json`; `cloudflared-felis.service` | the config inside every bundle; the login certificate, the tunnel's credentials and the unit are not copied | the Cloudflare step of `sudo felis setup` (step 11 below) | nothing: the tunnel, its DNS records and the Access application live at Cloudflare |
| Platform images, Velocity, the JRE, build tools | registry, `/opt/felis` | not copied | the installer builds and pushes them again | nothing |
| k3s itself (its token, CA, datastore, Secrets, Deployments) | `/var/lib/rancher/k3s` | not copied | the installer makes a new single-node cluster and renders every Secret and Deployment from `/etc/felis` | nothing: no Felis data lives only there |
@@ -2313,17 +2314,21 @@ and each skips what is already in place, so an interrupted one resumes. Write
those sizes down with the bucket's download rate and you have the recovery
time for your install. A whole-host rehearsal on a spare machine, once per
release, is the way to know it for sure. The spare follows the steps below
without step 8: it finds the production host named in the bucket and stands
by, so it copies nothing into the bucket, prunes nothing there and, while
production keeps writing, mails none of the owners its restored database
holds.
without steps 8 and 11. It finds the production host named in the bucket and
stands by, so it copies nothing into the bucket, prunes nothing there and,
while production keeps writing, mails none of the owners its restored
database holds. A second connector on the production tunnel would take a
share of the real visitors, so the spare leaves the tunnel alone and is
reached by its own address (step 10 moves it to one).
The order below matters: the state goes in before the installer so it reuses
the old secrets and bucket; the images go back before the database so the
servers the database restores find the digests they pin, inside the pruner's
24-hour grace; the MinecraftServers go back with the database that names their
owners; the worlds come last because a restore needs a server to restore
into.
owners; the worlds come after them because a restore needs a server to
restore into. Mail is checked before the domain moves, because a move can
cost passkeys and email codes are the way back in; the domain moves after the
MinecraftServers are applied, since the login gate's env carries it.
### Rebuild on a new host (the old one is gone)
@@ -2441,6 +2446,67 @@ host yourself, plus the off-site encryption key if the copy is in the bucket.
wrote it, and exits 4. With `-yes` it records this host; the old host, if it
ever runs again, copies nothing more and says it was taken over. Skip this
step on a rehearsal machine.
9. Check that this host can send mail, when the old one had a relay. Email
codes are how users sign in without a passkey, and step 10 can cost them
their passkeys. Run `sudo felis setup`, press `e` (configure email) on the
status screen and save the pre-filled relay with its password typed again
(the form never shows the stored one). Saving delivers one self-test
message to the From address and writes nothing unless it arrives. A relay
that admits senders by address, or an SPF record for the From domain that
lists the old host's address, refuses this host or sends its mail to spam:
add the new address there and save again. With no relay at all, an Owner
who cannot sign in recovers with `sudo felis breakGlass` (§17).
10. Move the install to this host's address when its names still lead to the
old one. The installer kept the bundle's root domain (its log says
`(reusing the installed domain)`).
- The `<old-address>.nip.io` default resolves to the dead host by
construction. Move it:
```
sudo felis domain set <new-address>.nip.io # the plan
sudo felis domain set -yes <new-address>.nip.io
sudo felis domain check
```
It runs after step 6 because the MinecraftServers applied there carry
the domain in the login gate's env. The plan counts the passkeys that
stop working; their users sign in with an email code (step 9) and
register a new one. docs/operations.md §6 has the rest of what it moves.
- A domain of your own stays. Point its records at the new address:
`<root>`, `*.<root>`, `console.<root>` and `op.console.<root>`, or only
`<root>` and `*.<root>` when the tunnel serves the panel. Confirm each
with `dig +short <name>` before going on: `felis domain check` warns
while a name does not resolve and accepts any address once it does.
Lower the records' TTL beforehand if the old host is still around to
plan with.
11. Bring the Cloudflare edge back, when the old host served the panel
through a tunnel. The bundle carries `/etc/felis/cloudflared.yml`; the
login certificate (`/root/.cloudflared/cert.pem`), the tunnel's
credentials (`/root/.cloudflared/<tunnel-id>.json`) and the
`cloudflared-felis` unit stay behind with the old disk, so the panel
hostnames answer Cloudflare error 1033 until a connector runs here.
Run `sudo felis setup`, press `c` (change connection) on the status
screen and choose Cloudflare Tunnel + Access. On the step's first screen
press `i` to install cloudflared, then `l` for `cloudflared tunnel login`
(browser consent on your account, which writes `cert.pem`); `enter` opens
the form once both are in place. Enter an API token and the account ID,
then the same Admit identity, hostnames and tunnel name as before. The
hostnames come pre-filled from the restored config and the tunnel name
defaults to `felis`; for another name, look up the id on the `tunnel:`
line of the restored `cloudflared.yml` in Zero Trust → Networks → Tunnels.
The step finds the existing tunnel by name and fetches its credentials
again, points the panel records at it, finds the existing Access
application (so its audience stays the one the restored config names)
and rewrites its policy, installs and starts `cloudflared-felis`, and
closes the panel's NodePort once the tunnel serves. A kept copy of the
old `<tunnel-id>.json` can go back into `/root/.cloudflared` (mode 0600)
first; the step still needs `cert.pem`. A different tunnel name makes a
second tunnel, moves the panel records to it and leaves the old one idle:
delete that one in the dashboard afterwards. Skip this step on a
rehearsal machine.
Check the rebuild before letting players in:
@@ -2449,13 +2515,24 @@ sudo felis db check # the database answers and has a fresh b
kubectl get minecraftservers -A # every server the bundle held
kubectl -n minecraft get pods # servers pull their pinned images (no ImagePullBackOff)
sudo felis offsite status # the hourly copy runs from this host again
sudo felis domain check # every name, the certificate and the tunnel on the new address
```
Point the panel and game hostnames at the new host (DNS, or the tunnel in
front of it). In the panel: sign in with an old account (accounts and passkeys
come back with the database), open a restored server's backup page and confirm
its archives are listed, restore the newest one, start the server and join
it.
A rehearsal machine that restored a tunnel install and moved to a `nip.io`
name in step 10 fails the Cloudflare tunnel line, since the tunnel still
routes production's names; that one failure is expected there.
In the panel, reached by its hostname (steps 10 and 11):
- Sign in with an old account; accounts and passkeys come back with the
database, except passkeys step 10 counted as lost.
- Sign out and sign in again with an email code, to an Owner account whose
inbox you read. The code arriving proves felis-api itself mails through the
relay to an outside inbox; step 9's self-test came from the setup console
and went to the From address. A code that never comes: §17.
- Open a restored server's backup page and confirm its archives are listed,
restore the newest one, start the server and join it at
`<name>.<root>`.
### Keep a copy somewhere else