Problem
claim_enable_jobs implements the lease acquisition correctly (atomic claim of NULL/expired with FOR UPDATE SKIP LOCKED, engram-postgres/src/lib.rs:2160-2192). But every post-claim write is unfenced — the classic lease-without-fencing-token bug: the lease arbitrates who starts, nothing arbitrates who may write.
update_enable_job_progress — WHERE id = $1 only, and it re-stamps claimed_at = NOW() (lib.rs:2194-2220, renewal at 2205) — extending the lease on behalf of whoever writes, including a pod whose lease already expired.
set_enable_job_state — WHERE id = $1 only (lib.rs:2222-2245).
record_enable_job_failure — WHERE id = $1 only, and it clears claimed_by/claimed_at unconditionally (lib.rs:2247-2269) — releasing a lease it may no longer hold.
- Only
retry_enable_job does CAS (AND state = 'failed', lib.rs:2277).
Failure scenario
- Pod A claims job J (chunking a 30 GiB image); GCS is slow; A blows past
lease_secs.
- Pod B's sweep legitimately re-claims J and restarts chunk-enable.
- A's stale ticks keep writing progress — each resets
chunks_done to A's smaller number and renews claimed_at, making B's claim look perpetually fresh-but-contested; the UI progress bar (GET /enable-jobs/:id) jumps backward.
- A hits a transient error →
record_enable_job_failure clears B's claim and stamps error on a job B is actively completing → a third pod claims it. Duplicate chunk uploads, attempts budget burned by phantom failures, jobs oscillating working→failed→pending while actually succeeding.
Proposed fix (mechanical)
- Thread the claimant identity through the trait:
update_enable_job_progress / set_enable_job_state / record_enable_job_failure in engram-core/src/traits/metadata.rs take claimant: &str.
- Add
AND claimed_by = $claimant to each UPDATE (lib.rs:2194-2269). On rows_affected() == 0 with the row existing, return MetaError::Conflict("lease lost to <claimed_by>") instead of success/NotFound.
- In the enable-job worker (caller in the coordinator's enable scanner), treat
Conflict as "stop work on this job immediately" — drop the local task without touching state.
- Keep the
claimed_at renewal in progress updates — it's now safe because it's fenced; document it as the lease heartbeat.
- Tests: extend
crates/engram-coordinator/tests/enable_jobs_live_pg.rs with a two-claimant scenario — claim as "pod-a", expire, claim as "pod-b", assert pod-a's progress/state/failure writes return Conflict and mutate nothing.
Acceptance criteria
Risk/scope
Small and mechanical — 3 SQL statements, 3 trait signatures, one worker branch, one test file. No schema change (claimed_by exists). 1-2 days.
Found by automated structural analysis (parallel codebase audit, 2026-06-12). Same fencing pattern as the session-lease release issue (#212) — consider one PR series establishing "every lease write is fenced" as a repo convention.
Problem
claim_enable_jobsimplements the lease acquisition correctly (atomic claim of NULL/expired withFOR UPDATE SKIP LOCKED,engram-postgres/src/lib.rs:2160-2192). But every post-claim write is unfenced — the classic lease-without-fencing-token bug: the lease arbitrates who starts, nothing arbitrates who may write.update_enable_job_progress—WHERE id = $1only, and it re-stampsclaimed_at = NOW()(lib.rs:2194-2220, renewal at 2205) — extending the lease on behalf of whoever writes, including a pod whose lease already expired.set_enable_job_state—WHERE id = $1only (lib.rs:2222-2245).record_enable_job_failure—WHERE id = $1only, and it clearsclaimed_by/claimed_atunconditionally (lib.rs:2247-2269) — releasing a lease it may no longer hold.retry_enable_jobdoes CAS (AND state = 'failed', lib.rs:2277).Failure scenario
lease_secs.chunks_doneto A's smaller number and renewsclaimed_at, making B's claim look perpetually fresh-but-contested; the UI progress bar (GET /enable-jobs/:id) jumps backward.record_enable_job_failureclears B's claim and stampserroron a job B is actively completing → a third pod claims it. Duplicate chunk uploads, attempts budget burned by phantom failures, jobs oscillatingworking→failed→pendingwhile actually succeeding.Proposed fix (mechanical)
update_enable_job_progress/set_enable_job_state/record_enable_job_failureinengram-core/src/traits/metadata.rstakeclaimant: &str.AND claimed_by = $claimantto each UPDATE (lib.rs:2194-2269). Onrows_affected() == 0with the row existing, returnMetaError::Conflict("lease lost to <claimed_by>")instead of success/NotFound.Conflictas "stop work on this job immediately" — drop the local task without touching state.claimed_atrenewal in progress updates — it's now safe because it's fenced; document it as the lease heartbeat.crates/engram-coordinator/tests/enable_jobs_live_pg.rswith a two-claimant scenario — claim as "pod-a", expire, claim as "pod-b", assert pod-a's progress/state/failure writes returnConflictand mutate nothing.Acceptance criteria
claimed_byfencing; zero-row UPDATE on an existing row surfaces asConflict.state,chunks_done,attempts, orclaimed_at.Conflict(mock-store unit test).Risk/scope
Small and mechanical — 3 SQL statements, 3 trait signatures, one worker branch, one test file. No schema change (
claimed_byexists). 1-2 days.Found by automated structural analysis (parallel codebase audit, 2026-06-12). Same fencing pattern as the session-lease release issue (#212) — consider one PR series establishing "every lease write is fenced" as a repo convention.