fix(quota): make the claim gate atomic, and stop zeroing the storage cache
Two defects in the §9.3 quota path, both invisible to the hermetic suite: - Audit #4's TOCTOU was real and documented: QuotaCheck and ClaimServer were separate statements, so two concurrent claims by one user for two different ownerless servers both read count < max_servers and both won. The gate now lives inside ClaimServer, in the SAME transaction as the ownership write, under pg_advisory_xact_lock(hashtext(user_id)) — the aggregate read, the four-dimension re-check (shared with QuotaCheck via one helper so the two cannot drift), and the UPDATE are one serialized decision. The loser gets ErrQuotaExceeded, which both claim handlers map to the same 403 the sequential path gives; the server row is additionally taken FOR UPDATE so same-server races still resolve to exactly one winner. - The server PATCH path called UpdateServerResources(..., 0) for storage even though a resources patch cannot change storage. The cached columns are the ONLY input to the quota aggregate, so every resource patch silently dropped that server's storage contribution from its owner's cap. The handler now reads the current spec and passes storage through. Red-then-green: the new pgint test drives two real concurrent claims against max_servers=1 (before: both win; now: exactly one win + one gated 403, and the DB shows one owned row); the hermetic suite pins the 403 mapping and the storage-preserving cache write.
This commit is contained in:
8 files changed
+274
-50
No files matched your search
@@ -142,6 +142,13 @@ func (a *API) handleClaim(w http.ResponseWriter, r *http.Request) {
|
||||
// ③ atomic claim
|
||||
claimed, err := a.Repo.ClaimServer(r.Context(), name, p.UserID)
|
||||
if err != nil {
|
||||
// The atomic gate re-checks quota under the per-user lock (audit #4): a
|
||||
// concurrent claim that spent the last slot surfaces here, with the same
|
||||
// 403 the pre-check gives sequentially.
|
||||
if errors.Is(err, ErrQuotaExceeded) {
|
||||
writeError(w, r, newError(http.StatusForbidden, "quota_exceeded", "server quota exhausted"))
|
||||
return
|
||||
}
|
||||
a.writeLookupError(w, r, err)
|
||||
return
|
||||
}
|
||||
@@ -759,7 +766,19 @@ func (a *API) handlePatchServer(w http.ResponseWriter, r *http.Request) {
|
||||
a.writeLookupError(w, r, err)
|
||||
return
|
||||
}
|
||||
_ = a.Repo.UpdateServerResources(r.Context(), name, newCPU, newMemMB, 0)
|
||||
// A resource patch cannot change storage, so its cached contribution must
|
||||
// be preserved: passing 0 would silently zero the storage dimension of the
|
||||
// owner's four-cap aggregate (the cached columns are its only input).
|
||||
storMB := 0
|
||||
if rec != nil {
|
||||
cur, err := a.Repo.ServerResources(r.Context(), name)
|
||||
if err != nil {
|
||||
writeError(w, r, err)
|
||||
return
|
||||
}
|
||||
storMB = cur.StorageMB
|
||||
}
|
||||
_ = a.Repo.UpdateServerResources(r.Context(), name, newCPU, newMemMB, storMB)
|
||||
} else {
|
||||
if err := a.Cluster.PatchServerSpec(r.Context(), name, patch); err != nil {
|
||||
a.writeLookupError(w, r, err)
|
||||
|
||||
Reference in new issue
Block a user