Files
Felis/cmd/felis/api.go
T
flyemoji fd062882ed feat(nano): give a Mojang player's name back to them, by prefixing the squatter
A premium player and a third-party player sharing a username could not both be
online. Whichever logged in second was kicked with "You are already connected to
this proxy!" -- even though the UUID rewrite had already made them two distinct
players on the backend. Velocity's player registry is keyed on the NAME (lowercased),
not the UUID, so two identities holding one name are one player as far as the proxy
is concerned, and the reclaim invariant the rewrite buys is invisible to it.

The fix needs no plugin and no state, because Velocity honours the name in the
hasJoined RESPONSE rather than pinning the one the client sent at login-start --
established by a real login, not by reading the source. So the multiplexer hands
back a different name and the collision is simply gone.

A third-party player whose name belongs to a Mojang account now joins as
PREFIX_name (LS_steve). Everyone else keeps their own name: the rename fires only
on an actual collision, decided by asking api.mojang.com whether the name is
registered. The name's owner is never the one renamed, which is 正版优先 falling out
for free -- the identity source is never rewritten, so there is no policy to encode
and no 30-day hold to track.

The premium-name answer is cached asymmetrically, because the two directions have
very different costs. "Taken" is nearly permanent (Mojang does not recycle names) and
is trusted for a day; "free" can stop being true the moment someone buys that name,
and a stale "free" leaves a squatter holding a name its real owner has just bought,
so it is trusted for ten minutes. A lookup that fails with nothing cached fails
CLOSED -- assume premium, rename the third-party player: a Mojang outage must not
become an opportunity to hold someone else's name, and being wrong that way costs a
cosmetic prefix while being wrong the other way bounces the name's owner off the
proxy. The lookup gets its own 2s client rather than sharing the 5s auth client,
since it is a SECOND Mojang round-trip on a login that already spent one.

prefix is a required, unique, 1-4 character config field rather than something
derived from the tag, because it is player-visible and no derivation can know that
"littleskin" is meant to read LS. Two sources sharing a prefix would rewrite their
same-named players onto one name, so uniqueness is enforced case-insensitively --
the proxy folds case, and LS/ls would collide there while reading as distinct here.

Also close a pre-existing hole on the path this touches: a third-party source's
profile name was relayed verbatim, so a hostile or sloppy Yggdrasil root could put
"§4admin", an empty string, or 200 characters straight into the proxy's player list.
The name is now checked against the Minecraft username charset and a bad one is a 204,
the same way a bad UUID already was.

Verified end to end on the deploy host (Velocity 3.5.1 + Paper 26.2), both branches:

  premium FLYEMOJ1     -> 195fadbd-f72e-4b9b-9f8f-f92586fe16ad, name unchanged
  LittleSkin FLYEMOJ1  -> LS_FLYEMOJ1, f1b7b6ae-f250-348a-b069-a2ec0fcae668
  both online at once, zero "already connected" rejections
  LittleSkin FelisNyaTest01 -> joins as FelisNyaTest01, no prefix, UUID still v3

The last line is the one that matters: an ordinary third-party player collides with
nobody and keeps their name, while the rewrite that keeps identities apart still ran.
Paper's "LS_FLYEMOJ1 (formerly known as li_FLYEMOJ1) joined the game" is the other
half of it -- the rename moved the player's display name and their playerdata came
along untouched, because every server-side key is the UUID and the UUID does not
depend on the name.

Known ceiling, left alone deliberately: two players of one source whose names agree
on their first 16-len(prefix)-1 characters truncate onto the same in-game name, and a
prefixed name may itself happen to be a premium name. Both cost an "already connected"
bounce, not an identity -- the UUID rewrite does not depend on the name at all.

BREAKING CHANGE: every [[auth_source]] now requires prefix = "XX" (1-4 letters or
digits, unique across sources). An existing nano felis.toml without it fails to load
with an error naming the field, rather than silently keeping the collision.
2026-07-13 13:00:54 +09:00

438 lines
20 KiB
Go

package main
import (
"context"
"flag"
"fmt"
"io"
"net/http"
"os"
"regexp"
"strings"
"time"
"felis.lolicon.best/internal/api"
"felis.lolicon.best/internal/apis/felis/v1alpha1"
"felis.lolicon.best/internal/backupjob"
"felis.lolicon.best/internal/build"
"felis.lolicon.best/internal/config"
"felis.lolicon.best/internal/panel"
"felis.lolicon.best/internal/passkey"
"felis.lolicon.best/internal/platform"
"felis.lolicon.best/internal/restore"
"felis.lolicon.best/internal/store"
"felis.lolicon.best/internal/submit"
"k8s.io/apimachinery/pkg/runtime"
utilruntime "k8s.io/apimachinery/pkg/util/runtime"
"k8s.io/client-go/kubernetes"
clientgoscheme "k8s.io/client-go/kubernetes/scheme"
ctrl "sigs.k8s.io/controller-runtime"
"sigs.k8s.io/controller-runtime/pkg/client"
)
// mojangSessionServer is the public Mojang hasJoined endpoint the Felis-nano multiplexer
// leads with as its code-owned identity anchor (正版优先). A protocol constant, not a
// deployment domain, so it is hardcoded rather than configured — and it is the ONLY source
// the code marks Identity (UUIDs trusted verbatim); config can never add another.
const mojangSessionServer = "https://sessionserver.mojang.com/session/minecraft/hasJoined"
// authSourcesFromConfig builds the multiplexer's priority list from the configured
// [[auth_source]] entries: Mojang leads as the code-owned identity anchor (正版优先, the ONLY
// Identity source — config can only append namespace-rewritten third-party sources, never a
// trusted one), then each configured source in file order. Both `felis api` and `felis nano`
// call it, so the "Mojang is prepended in code" invariant lives in exactly one place.
func authSourcesFromConfig(configured []config.AuthSourceConfig) []api.AuthSource {
sources := make([]api.AuthSource, 0, len(configured)+1)
sources = append(sources, api.AuthSource{Tag: "mojang", URL: mojangSessionServer, Identity: true})
for _, s := range configured {
sources = append(sources, api.AuthSource{Tag: s.Tag, Prefix: s.Prefix, URL: s.URL})
}
return sources
}
// cmdAPI runs felis-api: two listeners, two middleware chains (spec §7). The
// internal face (service token) is fully wired. The external face is wired but
// fails closed until an Access JWKS key function is configured — the verifier's
// audience logic is unit-tested (internal/api), the JWKS source is a deployment
// integration point.
func cmdAPI(args []string, stdout, stderr io.Writer) int {
fs := flag.NewFlagSet("api", flag.ContinueOnError)
fs.SetOutput(stderr)
cfgPath := fs.String("config", "/etc/felis/felis.toml", "path to felis.toml")
internalAddr := fs.String("internal-addr", ":8081", "internal-face listen address (service token, no Zero Trust)")
httpsAddr := fs.String("https-addr", "", "external HTTPS listen address (disabled unless --tls-cert and --tls-key are also set)")
tlsCert := fs.String("tls-cert", "", "TLS certificate path for --https-addr")
tlsKey := fs.String("tls-key", "", "TLS private key path for --https-addr")
if err := fs.Parse(args); err != nil {
return 2
}
if (*httpsAddr == "") != (*tlsCert == "" || *tlsKey == "") {
fmt.Fprintln(stderr, "felis api: --https-addr requires both --tls-cert and --tls-key")
return 2
}
cfg, err := config.Load(*cfgPath)
if err != nil {
fmt.Fprintf(stderr, "felis api: %v\n", err)
return 1
}
ctx := ctrl.SetupSignalHandler()
drv, err := store.Open(ctx, cfg.Database.URL)
if err != nil {
fmt.Fprintf(stderr, "felis api: open database: %v\n", err)
return 1
}
defer drv.Close()
scheme := runtime.NewScheme()
utilruntime.Must(clientgoscheme.AddToScheme(scheme))
utilruntime.Must(v1alpha1.AddToScheme(scheme))
// Both clients are built from the SAME rest.Config. The controller-runtime
// client.Client drives CRDs/Secrets/Jobs (cluster, console-write, restore); the
// typed clientset is needed solely for the read-side console, because the
// pods/log subresource (GetLogs(...).Stream) lives only on the typed CoreV1
// client, not on client.Client (spec §8 读=pods/log follow).
restCfg := ctrl.GetConfigOrDie()
cl, err := client.New(restCfg, client.Options{Scheme: scheme})
if err != nil {
fmt.Fprintf(stderr, "felis api: build k8s client: %v\n", err)
return 1
}
clientset, err := kubernetes.NewForConfig(restCfg)
if err != nil {
fmt.Fprintf(stderr, "felis api: build k8s clientset: %v\n", err)
return 1
}
token := os.Getenv("FELIS_SERVICE_TOKEN")
if token == "" {
fmt.Fprintln(stderr, "felis api: warning: FELIS_SERVICE_TOKEN unset — internal face will reject all callers")
}
// Build subsystem (spec §16): the weak-SA build Job runs in the configured
// build namespace and pushes to the internal registry. The build Pod never
// holds DB credentials — felis-api owns the PG store and admits scanned
// images, so the Builder is constructed here with both bindings.
builder := &build.Builder{
Store: build.NewPGStore(drv.DB()),
Jobs: build.NewK8sJobs(cl, buildConfig(cfg)),
Config: buildConfig(cfg),
}
// User-modpack approval lane (user-directed extension over §16; see
// internal/submit). An ordinary user may only SUBMIT a
// modpack; an admin must approve it before anything is built, at which point
// the SAME Trivy-gated Builder runs as for an admin's direct build. Registry
// MUST match the Builder's RegistryURL (cfg.Registry.URL) — both are wired from
// the one field here so the lane's pre-CAS validate and the Builder's Submit
// can never disagree about the push target.
//
// The blob upload transport is selected by the shape of user_uploads_context —
// the two backends the setup wizard chooses between. A local path wires
// LocalContextStore (the mounted uploads PVC); an s3:// base wires
// S3ContextStore when its credentials resolve. Either way the store's target is
// derived from the SAME config field the context ref uses, so the blob lands
// exactly where Kaniko's --context points. Anything else — or an s3:// base with
// no credentials configured — leaves Blobs nil so POST
// /me/submissions/{id}/context returns 503, honest like the restore executor
// when its PVC is not supplied. (Letting the sandboxed Kaniko build Pod READ the
// context — PVC mount for local, creds+egress for S3 — is a separate deployment
// integration.)
contextBase := cfg.Registry.UserUploadsContext
var blobs submit.Blobs
switch {
case isLocalUploadsPath(contextBase):
// Normalize a file:// URL to the plain path ONCE and feed it to BOTH the
// derived ref (ContextStore) and the store (Base), so the recorded
// context_ref and the on-disk write location can never diverge.
contextBase = strings.TrimPrefix(contextBase, "file://")
blobs = &submit.LocalContextStore{Base: contextBase}
case strings.HasPrefix(strings.ToLower(contextBase), "s3://"):
if s3, err := newS3UploadsStore(cfg.Registry); err != nil {
fmt.Fprintf(stderr, "felis api: S3 user-uploads store not configured (%v) — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", err)
} else {
blobs = s3
}
default:
fmt.Fprintf(stderr, "felis api: user-uploads context %q is neither a local path nor an s3:// base — modpack upload transport disabled (POST /api/v1/me/submissions/{id}/context returns 503)\n", contextBase)
}
submissions := &submit.Manager{
Store: submit.NewPGStore(drv.DB()),
Builds: builder,
Registry: cfg.Registry.URL,
ContextStore: contextBase,
Blobs: blobs,
}
// Restore subsystem (spec §7): the weak-SA restore Job mounts the target
// world PVC + the backup PVC and runs `felis restore`. It needs deployment-
// specific values that have no safe default — the felis image to run and the
// backup PVC to mount — so it is wired only when both are supplied. Otherwise
// the Restorer is left nil and the restore endpoint honestly returns 503
// rather than enqueuing a Job that cannot run. (The archive store no longer
// gates wiring here: config.Validate rejects any recognized-but-unimplemented
// store at load, so by this point cfg.Archive.Store is guaranteed tarLocal.)
var restorer api.Restorer
felisImage, backupPVC := os.Getenv("FELIS_IMAGE"), os.Getenv("FELIS_BACKUP_PVC")
if felisImage != "" && backupPVC != "" {
rcfg := restoreConfig(cfg, felisImage, backupPVC)
restorer = &restore.Restorer{Jobs: restore.NewK8sJobs(cl), Config: rcfg}
} else {
fmt.Fprintln(stderr, "felis api: restore executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — restore endpoint returns 503")
}
// On-demand backup subsystem (spec §18/§19 WorldArchiver, run on demand). Its
// backup Job mirrors the restore Job's weak-SA isolation but additionally mounts
// the config Secret so it self-records the world_backups row (see internal/
// backupjob). It needs the same deployment-specific values as restore, so it is
// wired under the same gate; otherwise the Backuper is left nil and the backup
// endpoint honestly returns 503.
var backuper api.Backuper
if felisImage != "" && backupPVC != "" {
backuper = &backupjob.Backuper{Jobs: backupjob.NewK8sJobs(cl), Config: backupConfig(cfg, felisImage, backupPVC)}
} else {
fmt.Fprintln(stderr, "felis api: backup executor disabled (needs FELIS_IMAGE and FELIS_BACKUP_PVC) — backup endpoint returns 503")
}
// One PGRepo instance backs both the handlers and the session verifier: the
// SessionAuth that fronts the external face reads sessions/users/settings from
// the same store the auth handlers write to, so a login and the next request
// agree on what local auth knows.
repo := api.NewPGRepo(drv.DB())
a := &api.API{
Repo: repo,
Cluster: api.NewK8sCluster(cl, cfg.K8s.Namespace),
Console: api.NewK8sConsole(cl, cfg.K8s.Namespace),
Logs: api.NewK8sLogStreamer(clientset, cfg.K8s.Namespace),
// Build-log stream (spec §16) is scoped to the BUILD namespace — the same
// value the Builder renders Jobs into — so it follows where build Pods run.
BuildLogs: api.NewK8sBuildLogStreamer(clientset, cfg.Registry.BuildNamespace),
Internal: api.BearerTokenAuth{Token: token},
Builder: builder,
Restorer: restorer,
Backuper: backuper,
Submissions: submissions,
// The external face is fronted by SessionAuth: it prefers a local-password
// session cookie and otherwise delegates to the Cloudflare-Access JWT verifier,
// so both auth models coexist on one face. The delegate's Keyfunc is
// intentionally nil — the JWT path fails closed until a JWKS-backed key function
// is wired (deployment integration point) — while the local-password path is
// live the moment `felis breakGlass` flips local_auth_enabled on.
External: api.SessionAuth{
Repo: repo,
Delegate: api.AccessVerifier{Audience: cfg.Auth.AccessJWTAud},
RootDomain: cfg.Server.RootDomain,
AdminHostname: cfg.Auth.AdminHostname,
},
RootDomain: cfg.Server.RootDomain,
WakeCooldown: 30 * time.Second,
// Bound concurrent console/build-log SSE streams per principal. Generous enough
// for legitimate multi-tab / multi-server watching, while capping how many
// upstream follow connections a single caller can tie up if their streams stall.
MaxStreamsPerPrincipal: 16,
}
fmt.Fprintln(stderr, "felis api: external face fails closed (Access JWKS key function not configured)")
// Felis-nano: wire the multi-source hasJoined multiplexer only when third-party auth
// sources are configured. Mojang leads as the code-owned identity anchor (正版优先);
// config can only append namespace-rewritten third-party sources, never a trusted one,
// so a misconfig cannot reopen the impersonation hole. No sources = a.AuthSources stays
// nil = the endpoint 204s every login (ships off).
if len(cfg.AuthSources) > 0 {
a.AuthSources = authSourcesFromConfig(cfg.AuthSources)
fmt.Fprintf(stderr, "felis api: hasJoined multiplexer active — Mojang + %d third-party source(s)\n", len(cfg.AuthSources))
}
// Passkey (WebAuthn) enrollment verifier (spec §14, Phase 6). The relying party is
// the panel (app) face: the RP id is the panel hostname and the single permitted
// origin is that host over https, so a credential enrolled here is scoped to the
// panel. It is wired only when auth.panel_hostname is configured; otherwise a.Passkey
// stays nil and the enrollment begin/finish routes honestly return 503 (the
// authenticated enrollment boundary is still enforced by the handlers). Scope is
// ENROLLMENT only — the login/assertion path is a deferred slice, and credentials
// enrolled under this RP id MUST be asserted under the same RP id when that slice
// lands. An admin passkey (if ever added) is a SEPARATE relying party on the admin
// host and is not wired here.
if cfg.Auth.PanelHostname != "" {
pv, err := passkey.New(cfg.Auth.PanelHostname, "Felis", []string{"https://" + cfg.Auth.PanelHostname})
if err != nil {
fmt.Fprintf(stderr, "felis api: passkey verifier disabled: %v — passkey endpoints return 503\n", err)
} else {
a.Passkey = pv
}
} else {
fmt.Fprintln(stderr, "felis api: passkey verifier disabled (auth.panel_hostname unset) — passkey endpoints return 503")
}
externalHandler := panel.Handler(a.ExternalHandler(), cfg.Server.RootDomain, cfg.Auth.PanelHostname, cfg.Auth.AdminHostname, resolvedVersion())
internalSrv := newAPIServer(*internalAddr, a.InternalHandler())
externalSrv := newAPIServer(cfg.Server.Listen, externalHandler)
errc := make(chan error, 3)
go func() { errc <- internalSrv.ListenAndServe() }()
go func() { errc <- externalSrv.ListenAndServe() }()
var httpsSrv *http.Server
if *httpsAddr != "" {
httpsSrv = newAPIServer(*httpsAddr, externalHandler)
go func() { errc <- httpsSrv.ListenAndServeTLS(*tlsCert, *tlsKey) }()
}
if httpsSrv != nil {
fmt.Fprintf(stdout, "felis api: internal=%s external=%s https=%s\n", *internalAddr, cfg.Server.Listen, *httpsAddr)
} else {
fmt.Fprintf(stdout, "felis api: internal=%s external=%s\n", *internalAddr, cfg.Server.Listen)
}
// reconcileBuilds drives the scan-gate translation: poll unfinished builds
// and advance any whose Job has reached a terminal phase. GET on a build also
// reconciles it, but this loop converges builds nobody is polling.
go reconcileBuilds(ctx, builder, stderr)
select {
case <-ctx.Done():
shutdownCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second)
defer cancel()
_ = internalSrv.Shutdown(shutdownCtx)
_ = externalSrv.Shutdown(shutdownCtx)
if httpsSrv != nil {
_ = httpsSrv.Shutdown(shutdownCtx)
}
return 0
case err := <-errc:
if err != nil && err != http.ErrServerClosed {
fmt.Fprintf(stderr, "felis api: listener exited: %v\n", err)
return 1
}
return 0
}
}
const (
// apiReadHeaderTimeout caps how long a client may take to send its request
// headers, defeating a Slowloris that trickles a header line forever to pin a
// connection open. It bounds only the header phase, so it is safe on every face —
// including the SSE streaming one, whose response, not its request, is long-lived.
apiReadHeaderTimeout = 10 * time.Second
// apiIdleTimeout caps how long a kept-alive connection may sit idle between
// requests before the server closes it, bounding idle-connection exhaustion.
apiIdleTimeout = 120 * time.Second
)
// newAPIServer builds an http.Server with hardened header/idle timeouts (gosec
// G112) shared by all three felis-api listeners (internal, external, https).
// WriteTimeout and ReadTimeout are deliberately LEFT UNSET: the external and https
// faces stream Server-Sent Events (console / build logs, spec §8) for the lifetime
// of a client's attachment, and a WriteTimeout would sever a healthy long-lived
// stream mid-flight. Slowloris is closed by ReadHeaderTimeout, which bounds only the
// header phase and never touches the response.
func newAPIServer(addr string, handler http.Handler) *http.Server {
return &http.Server{
Addr: addr,
Handler: handler,
ReadHeaderTimeout: apiReadHeaderTimeout,
IdleTimeout: apiIdleTimeout,
}
}
// buildConfig projects felis.toml onto the build subsystem config (spec §16,
// §24). Unset fields fall back to the build package's hardened defaults
// (felis-build namespace + weak SA, 30m deadline, resource limits).
func buildConfig(cfg *config.Config) build.Config {
return build.Config{
Namespace: cfg.Registry.BuildNamespace,
RegistryURL: cfg.Registry.URL,
}
}
// uploadsSchemeRE matches a leading URL scheme like "s3://" or "gs://".
var uploadsSchemeRE = regexp.MustCompile(`^[a-zA-Z][a-zA-Z0-9+.-]*://`)
// isLocalUploadsPath reports whether the user-uploads context base is a local
// filesystem path (a bare path or a file:// URL), i.e. one LocalContextStore can
// write to. An s3:// base routes to newS3UploadsStore instead; any other scheme
// has no implemented transport, so its uploads are left disabled (503).
func isLocalUploadsPath(base string) bool {
if strings.HasPrefix(base, "file://") {
return true
}
return !uploadsSchemeRE.MatchString(base)
}
// newS3UploadsStore builds the S3 blob transport for an s3:// user_uploads_context.
// The bucket + key prefix come from the base itself; the endpoint/region come from
// [registry.s3]; and the credentials are read from the environment variables named
// by access_key_ref / secret_key_ref (defaulting to the fixed env names the
// felis-api Deployment injects from the felis-uploads-s3 Secret). Any missing piece
// is an error, so the caller leaves Blobs nil and the upload endpoint returns 503
// rather than pretending it can persist a file.
func newS3UploadsStore(reg config.RegistryConfig) (submit.Blobs, error) {
accessRef, secretRef := reg.S3.AccessKeyRef, reg.S3.SecretKeyRef
if accessRef == "" {
accessRef = platform.UploadsS3AccessKeyEnv
}
if secretRef == "" {
secretRef = platform.UploadsS3SecretKeyEnv
}
accessKey, secretKey := os.Getenv(accessRef), os.Getenv(secretRef)
if accessKey == "" || secretKey == "" {
return nil, fmt.Errorf("credentials env %s/%s are empty", accessRef, secretRef)
}
return submit.NewS3ContextStore(submit.S3StoreConfig{
Base: reg.UserUploadsContext,
Endpoint: reg.S3.Endpoint,
Region: reg.S3.Region,
AccessKey: accessKey,
SecretKey: secretKey,
})
}
// restoreConfig projects felis.toml + the deployment-supplied image and backup
// PVC onto the restore subsystem config (spec §7). The runtime identity, mount
// roots, resource limits, and weak SA fall back to the restore package's
// hardened defaults. BackupRoot tracks cfg.Archive.LocalPath because tarLocal
// archive refs are absolute: the restore Pod must mount the backup PVC at the
// same path the reaper wrote archives under, or the stored ref won't resolve.
func restoreConfig(cfg *config.Config, image, backupPVC string) restore.Config {
return restore.Config{
Namespace: cfg.K8s.Namespace,
Image: image,
BackupPVC: backupPVC,
ArchiveStore: cfg.Archive.Store,
BackupRoot: cfg.Archive.LocalPath,
}
}
// backupConfig builds the on-demand backup executor's config from felis.toml plus
// the deployment-supplied image and backup PVC. BackupRoot mirrors restoreConfig —
// it MUST equal [archive] local_path so the recorded ref resolves the same way a
// later restore Job mounts it. ConfigSecret/ConfigMount are left to backupjob's
// defaults (the control-plane manifest names), which is the Secret this backup Job
// mounts to self-record its world_backups row.
func backupConfig(cfg *config.Config, image, backupPVC string) backupjob.Config {
return backupjob.Config{
Namespace: cfg.K8s.Namespace,
Image: image,
BackupPVC: backupPVC,
BackupRoot: cfg.Archive.LocalPath,
}
}
// reconcileBuilds polls unfinished builds on an interval and advances any whose
// Job has reached a terminal phase. It exits when ctx is cancelled.
func reconcileBuilds(ctx context.Context, b *build.Builder, stderr io.Writer) {
t := time.NewTicker(15 * time.Second)
defer t.Stop()
for {
select {
case <-ctx.Done():
return
case <-t.C:
if _, err := b.SyncAll(ctx); err != nil {
fmt.Fprintf(stderr, "felis api: build reconcile: %v\n", err)
}
}
}
}