Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.
felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.
What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.
Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.
The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.
Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:
* A write accepts arbitrary bytes at any path, so an owner can place a
loadable plugin jar. This is deliberate — it is what a hosting panel's
file manager does, scoped to a server the caller already owns and
already drives through /command — but it is the one owner-tier route
that lands executable code in a backend pod, since images are
admin-only and modpack submissions need an admin verdict.
* config/paper-global.yml is refused on read. felis-lobby's entrypoint
writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
identical on every backend, so reading it from a server you own would
hand you the handshake key for everyone else's. It is the only path in
the mount that is not the caller's own data, and therefore the only
denial. The comparison is on the cleaned path, or ./config/... would
walk straight through it.
Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.
The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
510 lines
23 KiB
Go
510 lines
23 KiB
Go
package main
|
|
|
|
import (
|
|
"context"
|
|
"flag"
|
|
"fmt"
|
|
"io"
|
|
"net/http"
|
|
"os"
|
|
"regexp"
|
|
"strings"
|
|
"time"
|
|
|
|
"felis.lolicon.best/internal/api"
|
|
"felis.lolicon.best/internal/apis/felis/v1alpha1"
|
|
"felis.lolicon.best/internal/backupjob"
|
|
"felis.lolicon.best/internal/build"
|
|
"felis.lolicon.best/internal/config"
|
|
"felis.lolicon.best/internal/fileedit"
|
|
"felis.lolicon.best/internal/mail"
|
|
"felis.lolicon.best/internal/panel"
|
|
"felis.lolicon.best/internal/passkey"
|
|
"felis.lolicon.best/internal/platform"
|
|
"felis.lolicon.best/internal/restore"
|
|
"felis.lolicon.best/internal/store"
|
|
"felis.lolicon.best/internal/submit"
|
|
"k8s.io/apimachinery/pkg/runtime"
|
|
utilruntime "k8s.io/apimachinery/pkg/util/runtime"
|
|
"k8s.io/client-go/kubernetes"
|
|
clientgoscheme "k8s.io/client-go/kubernetes/scheme"
|
|
ctrl "sigs.k8s.io/controller-runtime"
|
|
"sigs.k8s.io/controller-runtime/pkg/client"
|
|
)
|
|
|
|
// mojangSessionServer is the public Mojang hasJoined endpoint the Felis-nano multiplexer
|
|
// leads with as its code-owned identity anchor (正版优先). A protocol constant, not a
|
|
// deployment domain, so it is hardcoded rather than configured — and it is the ONLY source
|
|
// the code marks Identity (UUIDs trusted verbatim); config can never add another.
|
|
const mojangSessionServer = "https://sessionserver.mojang.com/session/minecraft/hasJoined"
|
|
|
|
// authSourcesFromConfig builds the multiplexer's priority list from the configured
|
|
// [[auth_source]] entries: Mojang leads as the code-owned identity anchor (正版优先, the ONLY
|
|
// Identity source — config can only append namespace-rewritten third-party sources, never a
|
|
// trusted one), then each configured source in file order. Both `felis api` and `felis nano`
|
|
// call it, so the "Mojang is prepended in code" invariant lives in exactly one place.
|
|
func authSourcesFromConfig(configured []config.AuthSourceConfig) []api.AuthSource {
|
|
sources := make([]api.AuthSource, 0, len(configured)+1)
|
|
sources = append(sources, api.AuthSource{Tag: "mojang", URL: mojangSessionServer, Identity: true})
|
|
for _, s := range configured {
|
|
sources = append(sources, api.AuthSource{Tag: s.Tag, Prefix: s.Prefix, URL: s.URL})
|
|
}
|
|
return sources
|
|
}
|
|
|
|
// cmdAPI runs felis-api: two listeners, two middleware chains (spec §7). The
|
|
// internal face (service token) is fully wired. The external face is wired but
|
|
// fails closed until an Access JWKS key function is configured — the verifier's
|
|
// audience logic is unit-tested (internal/api), the JWKS source is a deployment
|
|
// integration point.
|
|
func cmdAPI(args []string, stdout, stderr io.Writer) int {
|
|
fs := flag.NewFlagSet("api", flag.ContinueOnError)
|
|
fs.SetOutput(stderr)
|
|
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
|
|
internalAddr := fs.String("internal-addr", ":8081", "internal-face listen address (service token, no Zero Trust)")
|
|
httpsAddr := fs.String("https-addr", "", "external HTTPS listen address (disabled unless --tls-cert and --tls-key are also set)")
|
|
tlsCert := fs.String("tls-cert", "", "TLS certificate path for --https-addr")
|
|
tlsKey := fs.String("tls-key", "", "TLS private key path for --https-addr")
|
|
if err := fs.Parse(args); err != nil {
|
|
return 2
|
|
}
|
|
if (*httpsAddr == "") != (*tlsCert == "" || *tlsKey == "") {
|
|
fmt.Fprintln(stderr, "felis api: --https-addr requires both --tls-cert and --tls-key")
|
|
return 2
|
|
}
|
|
|
|
cfg, err := config.Load(*cfgPath)
|
|
if err != nil {
|
|
fmt.Fprintf(stderr, "felis api: %v\n", err)
|
|
return 1
|
|
}
|
|
|
|
ctx := ctrl.SetupSignalHandler()
|
|
|
|
drv, err := store.Open(ctx, cfg.Database.URL)
|
|
if err != nil {
|
|
fmt.Fprintf(stderr, "felis api: open database: %v\n", err)
|
|
return 1
|
|
}
|
|
defer drv.Close()
|
|
|
|
scheme := runtime.NewScheme()
|
|
utilruntime.Must(clientgoscheme.AddToScheme(scheme))
|
|
utilruntime.Must(v1alpha1.AddToScheme(scheme))
|
|
// Both clients are built from the SAME rest.Config. The controller-runtime
|
|
// client.Client drives CRDs/Secrets/Jobs (cluster, console-write, restore); the
|
|
// typed clientset is needed solely for the read-side console, because the
|
|
// pods/log subresource (GetLogs(...).Stream) lives only on the typed CoreV1
|
|
// client, not on client.Client (spec §8 读=pods/log follow).
|
|
restCfg := ctrl.GetConfigOrDie()
|
|
cl, err := client.New(restCfg, client.Options{Scheme: scheme})
|
|
if err != nil {
|
|
fmt.Fprintf(stderr, "felis api: build k8s client: %v\n", err)
|
|
return 1
|
|
}
|
|
clientset, err := kubernetes.NewForConfig(restCfg)
|
|
if err != nil {
|
|
fmt.Fprintf(stderr, "felis api: build k8s clientset: %v\n", err)
|
|
return 1
|
|
}
|
|
|
|
token := os.Getenv("FELIS_SERVICE_TOKEN")
|
|
if token == "" {
|
|
fmt.Fprintln(stderr, "felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers")
|
|
}
|
|
|
|
// Email one-time codes go through the [smtp] relay when one is configured; the
|
|
// password is read from the env var password_ref names (default SMTPPasswordEnv,
|
|
// injected from the felis-smtp Secret). No [smtp] host ⇒ mailer stays nil and
|
|
// deliverOTP logs each code server-side (the pre-SMTP bootstrap posture).
|
|
var mailer api.OTPMailer
|
|
if cfg.SMTP.Host != "" {
|
|
passRef := cfg.SMTP.PasswordRef
|
|
if passRef == "" {
|
|
passRef = platform.SMTPPasswordEnv
|
|
}
|
|
password := os.Getenv(passRef)
|
|
if cfg.SMTP.Username != "" && password == "" {
|
|
fmt.Fprintf(stderr, "felis api: warning: [smtp] username is set but credentials env %s is empty — OTP sends will fail AUTH\n", passRef)
|
|
}
|
|
mailer = &mail.SMTP{
|
|
Host: cfg.SMTP.Host,
|
|
Port: cfg.SMTP.Port,
|
|
From: cfg.SMTP.From,
|
|
Username: cfg.SMTP.Username,
|
|
Password: password,
|
|
}
|
|
} else {
|
|
fmt.Fprintln(stderr, "felis api: [smtp] not configured — email one-time codes are logged, not mailed")
|
|
}
|
|
|
|
// Build subsystem (spec §16): the weak-SA build Job runs in the configured
|
|
// build namespace and pushes to the internal registry. The build Pod never
|
|
// holds DB credentials — felis-api owns the PG store and admits scanned
|
|
// images, so the Builder is constructed here with both bindings.
|
|
builder := &build.Builder{
|
|
Store: build.NewPGStore(drv.DB()),
|
|
Jobs: build.NewK8sJobs(cl, buildConfig(cfg)),
|
|
Config: buildConfig(cfg),
|
|
}
|
|
|
|
// User-modpack approval lane (user-directed extension over §16; see
|
|
// internal/submit). An ordinary user may only SUBMIT a
|
|
// modpack; an admin must approve it before anything is built, at which point
|
|
// the SAME Trivy-gated Builder runs as for an admin's direct build. Registry
|
|
// MUST match the Builder's RegistryURL (cfg.Registry.URL) — both are wired from
|
|
// the one field here so the lane's pre-CAS validate and the Builder's Submit
|
|
// can never disagree about the push target.
|
|
//
|
|
// The blob upload transport is selected by the shape of user_uploads_context —
|
|
// the two backends the setup wizard chooses between. A local path wires
|
|
// LocalContextStore (the mounted uploads PVC); an s3:// base wires
|
|
// S3ContextStore when its credentials resolve. Either way the store's target is
|
|
// derived from the SAME config field the context ref uses, so the blob lands
|
|
// exactly where Kaniko's --context points. Anything else — or an s3:// base with
|
|
// no credentials configured — leaves Blobs nil so POST
|
|
// /me/submissions/{id}/context returns 503, honest like the restore executor
|
|
// when its PVC is not supplied. (Letting the sandboxed Kaniko build Pod READ the
|
|
// context — PVC mount for local, creds+egress for S3 — is a separate deployment
|
|
// integration.)
|
|
contextBase := cfg.Registry.UserUploadsContext
|
|
var blobs submit.Blobs
|
|
switch {
|
|
case isLocalUploadsPath(contextBase):
|
|
// Normalize a file:// URL to the plain path ONCE and feed it to BOTH the
|
|
// derived ref (ContextStore) and the store (Base), so the recorded
|
|
// context_ref and the on-disk write location can never diverge.
|
|
contextBase = strings.TrimPrefix(contextBase, "file://")
|
|
blobs = &submit.LocalContextStore{Base: contextBase}
|
|
case strings.HasPrefix(strings.ToLower(contextBase), "s3://"):
|
|
if s3, err := newS3UploadsStore(cfg.Registry); err != nil {
|
|
fmt.Fprintf(stderr, "felis api: S3 user-uploads store not configured (%v) — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", err)
|
|
} else {
|
|
blobs = s3
|
|
}
|
|
default:
|
|
fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase)
|
|
}
|
|
submissions := &submit.Manager{
|
|
Store: submit.NewPGStore(drv.DB()),
|
|
Builds: builder,
|
|
Registry: cfg.Registry.URL,
|
|
ContextStore: contextBase,
|
|
Blobs: blobs,
|
|
}
|
|
|
|
// Restore subsystem (spec §7): the weak-SA restore Job mounts the target
|
|
// world PVC + the backup PVC and runs `felis restore`. It needs deployment-
|
|
// specific values that have no safe default — the felis image to run and the
|
|
// backup PVC to mount — so it is wired only when both are supplied. Otherwise
|
|
// the Restorer is left nil and the restore endpoint honestly returns 503
|
|
// rather than enqueuing a Job that cannot run. (The archive store no longer
|
|
// gates wiring here: config.Validate rejects any recognized-but-unimplemented
|
|
// store at load, so by this point cfg.Archive.Store is guaranteed tarLocal.)
|
|
var restorer api.Restorer
|
|
felisImage, backupPVC := os.Getenv("FELIS_IMAGE"), os.Getenv("FELIS_BACKUP_PVC")
|
|
if felisImage != "" && backupPVC != "" {
|
|
rcfg := restoreConfig(cfg, felisImage, backupPVC)
|
|
restorer = &restore.Restorer{Jobs: restore.NewK8sJobs(cl), Config: rcfg}
|
|
} else {
|
|
fmt.Fprintln(stderr, "felis api: restore executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — restore endpoint returns 503")
|
|
}
|
|
|
|
// On-demand backup subsystem (spec §18/§19 WorldArchiver, run on demand). Its
|
|
// backup Job mirrors the restore Job's weak-SA isolation but additionally mounts
|
|
// the config Secret so it self-records the world_backups row (see internal/
|
|
// backupjob). It needs the same deployment-specific values as restore, so it is
|
|
// wired under the same gate; otherwise the Backuper is left nil and the backup
|
|
// endpoint honestly returns 503.
|
|
var backuper api.Backuper
|
|
if felisImage != "" && backupPVC != "" {
|
|
backuper = &backupjob.Backuper{Jobs: backupjob.NewK8sJobs(cl), Config: backupConfig(cfg, felisImage, backupPVC)}
|
|
} else {
|
|
fmt.Fprintln(stderr, "felis api: backup executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — backup endpoint returns 503")
|
|
}
|
|
|
|
// Server file editor: a weak-SA Job mounts ONLY the target world PVC and runs
|
|
// `felis files`, printing its result for felis-api to read back through
|
|
// pods/log (see internal/fileedit). It needs FELIS_IMAGE but — unlike restore
|
|
// and backup — no backup PVC, since it never touches the archive store, so it
|
|
// is wired on the image alone; otherwise the editor is left nil and the file
|
|
// endpoints honestly return 503. It takes the typed clientset rather than the
|
|
// controller-runtime client because the log subresource lives only on the typed
|
|
// CoreV1 client, and one client covers its Job create, Pod list, and log read.
|
|
var files api.FileEditor
|
|
if felisImage != "" {
|
|
files = &fileedit.Editor{
|
|
Runner: fileedit.NewK8sRunner(clientset),
|
|
Config: fileEditConfig(cfg, felisImage),
|
|
}
|
|
} else {
|
|
fmt.Fprintln(stderr, "felis api: file editor disabled (needs FELIS_IMAGE) — file endpoints return 503")
|
|
}
|
|
|
|
// One PGRepo instance backs both the handlers and the session verifier: the
|
|
// SessionAuth that fronts the external face reads sessions/users/settings from
|
|
// the same store the auth handlers write to, so a login and the next request
|
|
// agree on what local auth knows.
|
|
repo := api.NewPGRepo(drv.DB())
|
|
|
|
a := &api.API{
|
|
Repo: repo,
|
|
Cluster: api.NewK8sCluster(cl, cfg.K8s.Namespace),
|
|
Console: api.NewK8sConsole(cl, cfg.K8s.Namespace),
|
|
Logs: api.NewK8sLogStreamer(clientset, cfg.K8s.Namespace),
|
|
// Build-log stream (spec §16) is scoped to the BUILD namespace — the same
|
|
// value the Builder renders Jobs into — so it follows where build Pods run.
|
|
BuildLogs: api.NewK8sBuildLogStreamer(clientset, cfg.Registry.BuildNamespace),
|
|
Internal: api.BearerTokenAuth{Token: token},
|
|
Builder: builder,
|
|
Restorer: restorer,
|
|
Backuper: backuper,
|
|
Files: files,
|
|
Submissions: submissions,
|
|
Mailer: mailer,
|
|
// The external face is fronted by SessionAuth: it prefers a local session
|
|
// cookie (minted by the passwordless doors) and otherwise delegates to the
|
|
// Cloudflare-Access JWT verifier, so both auth models coexist on one face. The
|
|
// delegate's Keyfunc is intentionally nil — the JWT path fails closed until a
|
|
// JWKS-backed key function is wired (deployment integration point) — while the
|
|
// local session path is live the moment `felis breakGlass` flips
|
|
// local_auth_enabled on.
|
|
External: api.SessionAuth{
|
|
Repo: repo,
|
|
Delegate: api.AccessVerifier{Audience: cfg.Auth.AccessJWTAud},
|
|
RootDomain: cfg.Server.RootDomain,
|
|
AdminHostname: cfg.Auth.AdminHostname,
|
|
},
|
|
RootDomain: cfg.Server.RootDomain,
|
|
AdminHostname: cfg.Auth.AdminHostname,
|
|
PanelHostname: cfg.Auth.PanelHostname,
|
|
WakeCooldown: 30 * time.Second,
|
|
// Bound concurrent console/build-log SSE streams per principal. Generous enough
|
|
// for legitimate multi-tab / multi-server watching, while capping how many
|
|
// upstream follow connections a single caller can tie up if their streams stall.
|
|
MaxStreamsPerPrincipal: 16,
|
|
}
|
|
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
|
|
|
|
// Felis-nano: wire the multi-source hasJoined multiplexer only when third-party auth
|
|
// sources are configured. Mojang leads as the code-owned identity anchor (正版优先);
|
|
// config can only append namespace-rewritten third-party sources, never a trusted one,
|
|
// so a misconfig cannot reopen the impersonation hole. No sources = a.AuthSources stays
|
|
// nil = the endpoint 204s every login (ships off).
|
|
if len(cfg.AuthSources) > 0 {
|
|
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
|
|
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
|
|
}
|
|
|
|
// Passkey (WebAuthn) verifier (spec §14, Phase 6). One relying party spans BOTH
|
|
// web faces: the RP id is the panel hostname (console.<root>), and because that is
|
|
// a domain suffix of the operator host (op.console.<root>), a single credential
|
|
// enrolled once asserts on either face — one binding, usable on the player console
|
|
// AND the operator console. Both hosts are therefore listed as permitted origins,
|
|
// while the RP id stays the panel host so the credential's scope is ONE relying
|
|
// party, not two. Wired only when auth.panel_hostname is configured; otherwise
|
|
// a.Passkey stays nil and the passkey routes honestly return 503 (the authenticated
|
|
// enrollment boundary is still enforced by the handlers).
|
|
if cfg.Auth.PanelHostname != "" {
|
|
origins := []string{"https://" + cfg.Auth.PanelHostname}
|
|
if admin := defaultAdminHostname(cfg.Server.RootDomain, cfg.Auth.AdminHostname); admin != "" && admin != cfg.Auth.PanelHostname {
|
|
origins = append(origins, "https://"+admin)
|
|
}
|
|
pv, err := passkey.New(cfg.Auth.PanelHostname, "Felis", origins)
|
|
if err != nil {
|
|
fmt.Fprintf(stderr, "felis api: passkey verifier disabled: %v — passkey endpoints return 503\n", err)
|
|
} else {
|
|
a.Passkey = pv
|
|
}
|
|
} else {
|
|
fmt.Fprintln(stderr, "felis api: passkey verifier disabled (auth.panel_hostname unset) — passkey endpoints return 503")
|
|
}
|
|
|
|
// Derive the console hostnames when felis.toml leaves them unset, exactly as the
|
|
// setup/breakGlass paths do — otherwise the SPA cannot tell which face it is
|
|
// serving and falls back to the player console on op.console.<root>.
|
|
externalHandler := panel.Handler(a.ExternalHandler(), cfg.Server.RootDomain,
|
|
defaultPanelHostname(cfg.Server.RootDomain, cfg.Auth.PanelHostname),
|
|
defaultAdminHostname(cfg.Server.RootDomain, cfg.Auth.AdminHostname),
|
|
resolvedVersion())
|
|
internalSrv := newAPIServer(*internalAddr, a.InternalHandler())
|
|
externalSrv := newAPIServer(cfg.Server.Listen, externalHandler)
|
|
|
|
errc := make(chan error, 3)
|
|
go func() { errc <- internalSrv.ListenAndServe() }()
|
|
go func() { errc <- externalSrv.ListenAndServe() }()
|
|
var httpsSrv *http.Server
|
|
if *httpsAddr != "" {
|
|
httpsSrv = newAPIServer(*httpsAddr, externalHandler)
|
|
go func() { errc <- httpsSrv.ListenAndServeTLS(*tlsCert, *tlsKey) }()
|
|
}
|
|
if httpsSrv != nil {
|
|
fmt.Fprintf(stdout, "felis api: internal=%s external=%s https=%s\n", *internalAddr, cfg.Server.Listen, *httpsAddr)
|
|
} else {
|
|
fmt.Fprintf(stdout, "felis api: internal=%s external=%s\n", *internalAddr, cfg.Server.Listen)
|
|
}
|
|
|
|
// reconcileBuilds drives the scan-gate translation: poll unfinished builds
|
|
// and advance any whose Job has reached a terminal phase. GET on a build also
|
|
// reconciles it, but this loop converges builds nobody is polling.
|
|
go reconcileBuilds(ctx, builder, stderr)
|
|
|
|
select {
|
|
case <-ctx.Done():
|
|
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
|
|
defer cancel()
|
|
_ = internalSrv.Shutdown(shutdownCtx)
|
|
_ = externalSrv.Shutdown(shutdownCtx)
|
|
if httpsSrv != nil {
|
|
_ = httpsSrv.Shutdown(shutdownCtx)
|
|
}
|
|
return 0
|
|
case err := <-errc:
|
|
if err != nil && err != http.ErrServerClosed {
|
|
fmt.Fprintf(stderr, "felis api: listener exited: %v\n", err)
|
|
return 1
|
|
}
|
|
return 0
|
|
}
|
|
}
|
|
|
|
const (
|
|
// apiReadHeaderTimeout caps how long a client may take to send its request
|
|
// headers, defeating a Slowloris that trickles a header line forever to pin a
|
|
// connection open. It bounds only the header phase, so it is safe on every face —
|
|
// including the SSE streaming one, whose response, not its request, is long-lived.
|
|
apiReadHeaderTimeout = 10 * time.Second
|
|
// apiIdleTimeout caps how long a kept-alive connection may sit idle between
|
|
// requests before the server closes it, bounding idle-connection exhaustion.
|
|
apiIdleTimeout = 120 * time.Second
|
|
)
|
|
|
|
// newAPIServer builds an http.Server with hardened header/idle timeouts (gosec
|
|
// G112) shared by all three felis-api listeners (internal, external, https).
|
|
// WriteTimeout and ReadTimeout are deliberately LEFT UNSET: the external and https
|
|
// faces stream Server-Sent Events (console / build logs, spec §8) for the lifetime
|
|
// of a client's attachment, and a WriteTimeout would sever a healthy long-lived
|
|
// stream mid-flight. Slowloris is closed by ReadHeaderTimeout, which bounds only the
|
|
// header phase and never touches the response.
|
|
func newAPIServer(addr string, handler http.Handler) *http.Server {
|
|
return &http.Server{
|
|
Addr: addr,
|
|
Handler: handler,
|
|
ReadHeaderTimeout: apiReadHeaderTimeout,
|
|
IdleTimeout: apiIdleTimeout,
|
|
}
|
|
}
|
|
|
|
// buildConfig projects felis.toml onto the build subsystem config (spec §16,
|
|
// §24). Unset fields fall back to the build package's hardened defaults
|
|
// (felis-build namespace + weak SA, 30m deadline, resource limits).
|
|
func buildConfig(cfg *config.Config) build.Config {
|
|
return build.Config{
|
|
Namespace: cfg.Registry.BuildNamespace,
|
|
RegistryURL: cfg.Registry.URL,
|
|
}
|
|
}
|
|
|
|
// uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://".
|
|
var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`)
|
|
|
|
// isLocalUploadsPath reports whether the user-uploads context base is a local
|
|
// filesystem path (a bare path or a file:// URL), i.e. one LocalContextStore can
|
|
// write to. An s3:// base routes to newS3UploadsStore instead; any other scheme
|
|
// has no implemented transport, so its uploads are left disabled (503).
|
|
func isLocalUploadsPath(base string) bool {
|
|
if strings.HasPrefix(base, "file://") {
|
|
return true
|
|
}
|
|
return !uploadsSchemeRE.MatchString(base)
|
|
}
|
|
|
|
// newS3UploadsStore builds the S3 blob transport for an s3:// user_uploads_context.
|
|
// The bucket + key prefix come from the base itself; the endpoint/region come from
|
|
// [registry.s3]; and the credentials are read from the environment variables named
|
|
// by access_key_ref / secret_key_ref (defaulting to the fixed env names the
|
|
// felis-api Deployment injects from the felis-uploads-s3 Secret). Any missing piece
|
|
// is an error, so the caller leaves Blobs nil and the upload endpoint returns 503
|
|
// rather than pretending it can persist a file.
|
|
func newS3UploadsStore(reg config.RegistryConfig) (submit.Blobs, error) {
|
|
accessRef, secretRef := reg.S3.AccessKeyRef, reg.S3.SecretKeyRef
|
|
if accessRef == "" {
|
|
accessRef = platform.UploadsS3AccessKeyEnv
|
|
}
|
|
if secretRef == "" {
|
|
secretRef = platform.UploadsS3SecretKeyEnv
|
|
}
|
|
accessKey, secretKey := os.Getenv(accessRef), os.Getenv(secretRef)
|
|
if accessKey == "" || secretKey == "" {
|
|
return nil, fmt.Errorf("credentials env %s/%s are empty", accessRef, secretRef)
|
|
}
|
|
return submit.NewS3ContextStore(submit.S3StoreConfig{
|
|
Base: reg.UserUploadsContext,
|
|
Endpoint: reg.S3.Endpoint,
|
|
Region: reg.S3.Region,
|
|
AccessKey: accessKey,
|
|
SecretKey: secretKey,
|
|
})
|
|
}
|
|
|
|
// restoreConfig projects felis.toml + the deployment-supplied image and backup
|
|
// PVC onto the restore subsystem config (spec §7). The runtime identity, mount
|
|
// roots, resource limits, and weak SA fall back to the restore package's
|
|
// hardened defaults. BackupRoot tracks cfg.Archive.LocalPath because tarLocal
|
|
// archive refs are absolute: the restore Pod must mount the backup PVC at the
|
|
// same path the reaper wrote archives under, or the stored ref won't resolve.
|
|
func restoreConfig(cfg *config.Config, image, backupPVC string) restore.Config {
|
|
return restore.Config{
|
|
Namespace: cfg.K8s.Namespace,
|
|
Image: image,
|
|
BackupPVC: backupPVC,
|
|
ArchiveStore: cfg.Archive.Store,
|
|
BackupRoot: cfg.Archive.LocalPath,
|
|
}
|
|
}
|
|
|
|
// backupConfig builds the on-demand backup executor's config from felis.toml plus
|
|
// the deployment-supplied image and backup PVC. BackupRoot mirrors restoreConfig —
|
|
// it MUST equal [archive] local_path so the recorded ref resolves the same way a
|
|
// later restore Job mounts it. ConfigSecret/ConfigMount are left to backupjob's
|
|
// defaults (the control-plane manifest names), which is the Secret this backup Job
|
|
// mounts to self-record its world_backups row.
|
|
func backupConfig(cfg *config.Config, image, backupPVC string) backupjob.Config {
|
|
return backupjob.Config{
|
|
Namespace: cfg.K8s.Namespace,
|
|
Image: image,
|
|
BackupPVC: backupPVC,
|
|
BackupRoot: cfg.Archive.LocalPath,
|
|
}
|
|
}
|
|
|
|
// fileEditConfig builds the file editor's config from felis.toml plus the
|
|
// deployment-supplied image. It is the shortest of the three: the editor mounts
|
|
// only the world PVC, so it needs no archive coordinates at all, and everything
|
|
// else — the weak SA, the "/data" world root that makes paths match what the
|
|
// minecraft server itself sees, the runtime identity, and the size/time ceilings —
|
|
// falls back to the fileedit package's hardened defaults.
|
|
func fileEditConfig(cfg *config.Config, image string) fileedit.Config {
|
|
return fileedit.Config{
|
|
Namespace: cfg.K8s.Namespace,
|
|
Image: image,
|
|
}
|
|
}
|
|
|
|
// reconcileBuilds polls unfinished builds on an interval and advances any whose
|
|
// Job has reached a terminal phase. It exits when ctx is cancelled.
|
|
func reconcileBuilds(ctx context.Context, b *build.Builder, stderr io.Writer) {
|
|
t := time.NewTicker(15 * time.Second)
|
|
defer t.Stop()
|
|
for {
|
|
select {
|
|
case <-ctx.Done():
|
|
return
|
|
case <-t.C:
|
|
if _, err := b.SyncAll(ctx); err != nil {
|
|
fmt.Fprintf(stderr, "felis api: build reconcile: %v\n", err)
|
|
}
|
|
}
|
|
}
|
|
}
|