Give an owner a way to repair the one failure no other endpoint covers: a
server that will not boot because a single line of server.properties or a
plugin's YAML is wrong. Until now that needed a human with cluster access.
felis-api cannot touch a world in-process — the world PVC is ReadWriteOnce
and its lifecycle belongs to the operator's StatefulSet — so the work runs
as a one-shot Job, and the server must be stopped first because a running
one holds the volume. That is the same constraint that shapes restore and
backup, and the handlers enforce the stopped gate the same way.
What is different is that the caller wants the OUTPUT, not just the side
effect. The Job prints its result to stdout and felis-api reads it back
through the pods/log subresource, which needs no permission felis-api does
not already hold: jobs:create, pods:list, pods/log:get. No pods/exec, no
pods/portforward, not even pods:get. The price is latency — every operation
is a Pod schedule — which is why this is a repair tool and not a file
manager.
Containment is structural, not textual. Every filesystem access goes through
os.Root, the stdlib's escape-proof directory handle, which resolves each
component against the open root descriptor and refuses "..", absolute paths,
and symlinks leading outside. The string-prefix check used elsewhere is not
reused here: it validates a path as text and then opens it as a path, and a
world directory holds attacker-influenced content, so a symlink swapped in
between those two steps is a live threat rather than a theoretical one.
os.Root has no such window because the check and the open are one operation.
The Job's isolation is a strict subset of a restore Pod's: the weak
felis-restore SA with its token auto-mount disabled, exactly one volume (the
world PVC, mounted read-only for list and read so two of the three
operations cannot mutate anything), no Secret, no ConfigMap, no database
URL, non-root with an fsGroup matching the operator's so a written file is
readable by the server that later mounts it, and backoffLimit 0 so a failed
write is never silently retried as a second write.
Two limits on the surface are worth stating plainly, because the mount is
the server's whole working directory rather than a config subtree:
* A write accepts arbitrary bytes at any path, so an owner can place a
loadable plugin jar. This is deliberate — it is what a hosting panel's
file manager does, scoped to a server the caller already owns and
already drives through /command — but it is the one owner-tier route
that lands executable code in a backend pod, since images are
admin-only and modpack submissions need an admin verdict.
* config/paper-global.yml is refused on read. felis-lobby's entrypoint
writes FELIS_FORWARDING_SECRET into it on every boot, and that value is
identical on every backend, so reading it from a server you own would
hand you the handshake key for everyone else's. It is the only path in
the mount that is not the caller's own data, and therefore the only
denial. The comparison is on the cleaned path, or ./config/... would
walk straight through it.
Writing that file is still allowed: it leaks nothing, and the entrypoint
rewrites it whole on every boot regardless.
The write body's content field is a *[]byte rather than a []byte for the
reason permissionRequest.Value is a *bool — a plain slice makes absent,
null, and empty indistinguishable, so a body of {} would decode to nil and
truncate the target to zero bytes while answering 200, destroying the very
config the caller opened the editor to repair.
188 lines
7.2 KiB
Go
188 lines
7.2 KiB
Go
package fileedit
|
|
|
|
import (
|
|
"bufio"
|
|
"context"
|
|
"errors"
|
|
"fmt"
|
|
"io"
|
|
"strings"
|
|
"time"
|
|
|
|
corev1 "k8s.io/api/core/v1"
|
|
metav1 "k8s.io/apimachinery/pkg/apis/meta/v1"
|
|
"k8s.io/client-go/kubernetes"
|
|
)
|
|
|
|
// pollInterval is how often the runner re-Lists Pods while waiting for the file
|
|
// Job to finish. It is a POLL rather than a Watch because felis-api holds
|
|
// pods:list and NOT pods:watch (internal/platform.APIMinecraftRole) — establishing
|
|
// a watch would need a permission this design exists to avoid. Half a second is
|
|
// well inside the human-perceptible floor for an operation already dominated by
|
|
// Pod scheduling, while keeping the request count on a slow image pull modest.
|
|
const pollInterval = 500 * time.Millisecond
|
|
|
|
// maxLogBytes bounds what the runner will buffer from a Pod's log. The payload is
|
|
// at most a base64-encoded MaxReadBytes (≈4/3 of 1 MiB) plus the JSON envelope, so
|
|
// 4 MiB is generous headroom while still refusing to let a Pod that floods stderr
|
|
// pull felis-api's memory down with it.
|
|
const maxLogBytes = 4 << 20
|
|
|
|
// K8sRunner is the production Runner (spec §7, §16). It drives one file operation
|
|
// end to end using ONLY the three permissions felis-api already holds in the
|
|
// minecraft namespace, which is the entire point of the design:
|
|
//
|
|
// jobs:create → create the file Job
|
|
// pods:list → find its Pod and observe the phase (no pods:get, no pods:watch)
|
|
// pods/log:get → read the printed result back
|
|
//
|
|
// It deliberately takes a typed kubernetes.Interface rather than the
|
|
// controller-runtime client that internal/restore's K8sJobs uses: the log
|
|
// subresource (GetLogs(...).Stream) exists only on the typed CoreV1 client, and
|
|
// the Job create is available on both — so one client covers all three calls
|
|
// instead of the binding carrying two.
|
|
//
|
|
// INTEGRATION-ONLY: like K8sLogStreamer and K8sCluster this needs a live cluster;
|
|
// it compiles here but is exercised only against one, never by the hermetic test
|
|
// suite. The Oracle verifies the layer above it (Editor orchestration and error
|
|
// mapping) against a fake Runner, and the Job shape via the pure jobspec.
|
|
type K8sRunner struct {
|
|
cs kubernetes.Interface
|
|
}
|
|
|
|
// NewK8sRunner builds a Runner over cs. Every per-operation parameter — the
|
|
// namespace included — travels in the JobParams the Editor renders, so there is no
|
|
// Config to retain here.
|
|
func NewK8sRunner(cs kubernetes.Interface) *K8sRunner {
|
|
return &K8sRunner{cs: cs}
|
|
}
|
|
|
|
// Run creates the file Job, waits for its Pod to reach a terminal phase, and
|
|
// returns the JSON payload from the ResultPrefix line of that Pod's log.
|
|
//
|
|
// The Job name carries a fresh random OpID (FilesJobName), so a create collision is
|
|
// not an expected condition the way it is for restore — an AlreadyExists here means
|
|
// a 64-bit collision inside one TTL window and is reported rather than absorbed,
|
|
// because absorbing it would mean returning ANOTHER operation's output.
|
|
func (k *K8sRunner) Run(ctx context.Context, p JobParams) ([]byte, error) {
|
|
job, err := FilesJob(p)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
if _, err := k.cs.BatchV1().Jobs(p.Namespace).Create(ctx, job, metav1.CreateOptions{}); err != nil {
|
|
return nil, fmt.Errorf("fileedit: create file job: %w", err)
|
|
}
|
|
|
|
pod, err := k.awaitPod(ctx, p)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
|
|
log, err := k.podLog(ctx, p.Namespace, pod.Name)
|
|
if err != nil {
|
|
return nil, err
|
|
}
|
|
|
|
payload, ok := extractResult(log)
|
|
if !ok {
|
|
// No marked line: the entrypoint died before printing (an unmountable volume,
|
|
// an OOM kill, a deadline). The log tail travels in the error for the operator's
|
|
// benefit — this error reaches felis-api's logs, while the caller gets the
|
|
// generic 500 writeError produces, so no node detail leaks to the browser.
|
|
return nil, fmt.Errorf("fileedit: file job %s produced no result (phase %s): %s",
|
|
job.Name, pod.Status.Phase, tail(log))
|
|
}
|
|
return payload, nil
|
|
}
|
|
|
|
// awaitPod polls until the operation's Pod reaches a terminal phase. It selects by
|
|
// the per-invocation LabelOpID, so it can never observe a different operation's Pod
|
|
// — the reason that label exists.
|
|
//
|
|
// Both Succeeded and Failed are terminal and BOTH return the Pod rather than an
|
|
// error, because a caller-fault result (a path that escapes the root, a file that
|
|
// is too large) is printed and then exited on cleanly, and even a genuinely failed
|
|
// Pod may have printed a diagnosable result first. Deciding what the outcome MEANS
|
|
// is the caller's job (Run reads the printed result); this function only decides
|
|
// when there is nothing left to wait for.
|
|
func (k *K8sRunner) awaitPod(ctx context.Context, p JobParams) (*corev1.Pod, error) {
|
|
ticker := time.NewTicker(pollInterval)
|
|
defer ticker.Stop()
|
|
|
|
for {
|
|
pods, err := k.cs.CoreV1().Pods(p.Namespace).List(ctx, metav1.ListOptions{
|
|
LabelSelector: LabelOpID + "=" + p.OpID,
|
|
})
|
|
if err != nil {
|
|
return nil, fmt.Errorf("fileedit: find file job pod: %w", err)
|
|
}
|
|
for i := range pods.Items {
|
|
switch pods.Items[i].Status.Phase {
|
|
case corev1.PodSucceeded, corev1.PodFailed:
|
|
return &pods.Items[i], nil
|
|
}
|
|
}
|
|
|
|
select {
|
|
case <-ticker.C:
|
|
case <-ctx.Done():
|
|
// The Editor's Timeout (or the client disconnecting) fired. The Job is left
|
|
// alone deliberately: felis-api holds no jobs:delete, and the Job's own
|
|
// activeDeadlineSeconds plus ttlSecondsAfterFinished retire it without help.
|
|
return nil, fmt.Errorf("fileedit: timed out waiting for the file job to finish: %w", ctx.Err())
|
|
}
|
|
}
|
|
}
|
|
|
|
// podLog reads a finished Pod's log. Follow is off — the Pod has already
|
|
// terminated, so the log is complete and a follow would merely block until the
|
|
// stream closed.
|
|
func (k *K8sRunner) podLog(ctx context.Context, namespace, pod string) (string, error) {
|
|
stream, err := k.cs.CoreV1().Pods(namespace).GetLogs(pod, &corev1.PodLogOptions{
|
|
Container: containerName,
|
|
}).Stream(ctx)
|
|
if err != nil {
|
|
return "", fmt.Errorf("fileedit: read file job log: %w", err)
|
|
}
|
|
defer stream.Close()
|
|
|
|
b, err := io.ReadAll(io.LimitReader(stream, maxLogBytes))
|
|
if err != nil && !errors.Is(err, io.EOF) {
|
|
return "", fmt.Errorf("fileedit: read file job log: %w", err)
|
|
}
|
|
return string(b), nil
|
|
}
|
|
|
|
// extractResult finds the marked payload in a Pod log. It scans for the LAST line
|
|
// carrying ResultPrefix because pods/log returns stdout and stderr MERGED: a Go
|
|
// runtime warning or a libc message can appear anywhere in the stream, so the
|
|
// payload must be located by its marker rather than by position. Taking the last
|
|
// match rather than the first is the conservative choice — if a marker somehow
|
|
// appeared more than once, the final one is the operation's actual outcome.
|
|
func extractResult(log string) ([]byte, bool) {
|
|
var payload string
|
|
var found bool
|
|
sc := bufio.NewScanner(strings.NewReader(log))
|
|
sc.Buffer(make([]byte, 0, 64*1024), maxLogBytes)
|
|
for sc.Scan() {
|
|
if rest, ok := strings.CutPrefix(sc.Text(), ResultPrefix); ok {
|
|
payload, found = rest, true
|
|
}
|
|
}
|
|
if !found {
|
|
return nil, false
|
|
}
|
|
return []byte(payload), true
|
|
}
|
|
|
|
// tail returns the last few hundred bytes of a log for an error message, so a
|
|
// diagnostic is useful without embedding an entire log in an error string.
|
|
func tail(log string) string {
|
|
const n = 512
|
|
log = strings.TrimSpace(log)
|
|
if len(log) <= n {
|
|
return log
|
|
}
|
|
return "..." + log[len(log)-n:]
|
|
}
|