Repository navigation
R2 Track A: effect queue + Router-driven workload + the expected-state model oracle - #786
Merged
Merged
Conversation
…faults) Pre-R2 every SimHostClient verb mutated world truth inline after one optional delay, so RPC loss/reorder/duplication and a replica crash BETWEEN the store commit (the ack the coordinator already holds) and the host-side world effect were structurally impossible (the audit's finding #2/#4). A mutating verb now records its world mutation as a typed `Effect`. Inline delivery is the default — byte-for-byte the pre-R2 world, so Calm seeds are unchanged (Calm never opens a deferred window). When the target host is in the deferred set (a Chaos fault window), the effect is queued under a monotonic serial (BTreeMap → deterministic order) for a later scheduler step to deliver / drop / duplicate / reorder: - DeferHost(i, on) — open/close a host's deferred window - DeliverEffects — deliver the queue in serial (causal) order - DropEffect — loss: drop one queued effect (seeded pick) - DuplicateEffect — re-deliver one effect (seeded pick) - ReorderEffects — deliver in a seeded-shuffled order Crash/restart of a host severs its in-flight effects (never resurrected onto the cleared VM set); quiescence closes every window and flushes the queue in order before the fleet heals, so world truth is consistent with the coordinator's commits. The Chaos weight table carves 9 points out of AdvanceTime/Driver/HostHeartbeats/Crash/RestartHost for the new arms, which shifts every Chaos seed's exploration (seeds pin to a commit); the pinned chaos seeds still converge and are kept. tests/effect_queue.rs pins the two reachable states the queue unlocks: a deferred create withheld-then-delivered, and a replica crash inside the commit→effect window followed by effect loss, both converging at quiescence (no_op_dropped + no-stragglers + no-orphans). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Pre-R2 the workload called store/core fns directly (0/17 HTTP, 0/~69 gRPC driven — the audit's finding). This lands the real surface: - workload.rs drives each replica's ACTUAL surface: the tonic AppSessionService/AppFleetService impls (auth + convert.rs + the same *_core the axum handler calls — handler-direct, no socket, to keep the paused-clock current-thread determinism) for create/prompt/resume/ delete/drain, and the actual axum api::router via tower::oneshot (real-wire, middleware + extractors) for admin pause/resume and the harness-idle host ingestion. A module honesty table documents which verb takes which path. - tests/api_surface.rs exercises it deterministically: create is acked, persisted (never lost), and REPLAYS byte-identically; prompt+delete round-trip; the axum router runs the real HTTP handlers to a normal response. Real bug found + fixed: create_session's prepare path minted the session id via a raw `SessionId::new()` (Uuid::new_v4) instead of the injected `services.entropy` — an ADR 0098 D1 determinism leak (the id diverged every replay; in prod OsEntropy makes this behavior-identical). The API-create path now replays. SimMeta gains four PG-faithful methods the real handlers reach (previously panic-stubs): insert/get/delete_broker_token, get/set_teleport_target, rebind_session_guarded, list_active_assignments_with_budgets_on_host — with broker-token and teleport-target conformance scenarios (ADR 0098 D4). Scope note: folding this workload into the chaos/calm SWARM additionally needs host-fidelity fixes (schedulable wire_version, staged ready_images) that unmask a genuine, separate driver-double-boot + eviction-durability class the pre-R2 sim silently suppressed by leaving hosts wire-skewed and snapshots non-recoverable. Those are real findings for a follow-up; here the surface is driven directly and deterministically rather than shipping a red swarm lane. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TigerBeetle's auditor shape: an expected-state model fed ONLY by ACKED workload outcomes, diffed against world/SimMeta truth in the standing invariant pass (every step + at quiescence). The honesty boundary is explicit — because the model is fed only by acks, a lost-response op (a create whose boot never durably established, an op whose reply dropped) is legitimately absent from the model and never asserted on. What it asserts: - acked-create/resume never silently lost: a session the workload saw reach Active (durable row committed) still has a row on every later step, unless a later ACKED destroy retired it. A row vanishing under a live session is exactly the #570 symptom class (coordinator-unbind vs host-teardown-reconcile racing a session out of existence) and the durability-lie class this program exists to catch. - read-your-acked-writes: the image a session was created with is never repainted (both replicas read the same shared SimMeta, so a present row is readable on either). Fed from the op-path CreateSession/ResumeSession outcomes (record_if_live records a session once observed Active). Wired into run()'s per-step and quiescence checks alongside the existing oracles. Non-vacuity is proven, not assumed: tests/model_oracle.rs drives a session to Active, drops its row directly (the corruption the #570 race produces), and asserts the auditor FIRES with `model-acked-create-not- lost`; restoring the row clears it. `SimMetadataStore::with_db_mut` is the test-only corruption injector. Swarm (chaos 0..40 x600) stays green — the auditor is a safety invariant the current code upholds; the canary test is what proves it can bite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
Author
|
Orchestrator review: the sessions.rs entropy fix is behavior-identical in prod (OsEntropy) and closes a live D1 leak in a driven path — the exact hole the R1.7b ids.rs proposal predicted. The scope call (dedicated deterministic API tests now; swarm integration held until #787's double-boot is fixed) is the right honesty tradeoff — a red lane teaches nothing. The double-boot finding is filed as #787 with your reproduction notes and marked as the next wave's opener. Merging on green. 🤖 Generated with Claude Code |
nikhilunni
marked this pull request as ready for review
July 18, 2026 21:29
This was referenced Jul 18, 2026
Closed
nikhilunni
added a commit
that referenced
this pull request
Jul 19, 2026
… assertions ADR 0098 Phase 3 wave 4 — the deferred workload-verb fold-in from #786/#793. Folds the deterministic API verbs #786 wired (but drove only in dedicated tests) into BOTH swarm profiles at small weights, each over the real service surface and each feeding the acked-only expected-state model oracle: - Prompt: the full send_prompt (gRPC) -> real Deliver op -> run_started ack loop. Active-only + the ack closes the loop so an unacked outbox row cannot redeliver forever (no real guest harness emits the ack in the sim). - Rename: the coordinator-owned `set_session_suggested_title` write (the exact call the harness-title sink performs), store-direct — the real-wire harness route's event-bus publish races the paused clock under time advances. - Destroy: the real Destroy op + teardown, driven through the op pipeline like ResumeSession (the client-facing delete_session wait-loop can't step under the single-step-per-pick sim model; its full handler stays in api_surface). The model oracle grows two assertions (non-vacuity proven in tests/workload_verbs.rs): - acked-destroy-never-resurrects: a torn-down session that reappears live is the double-boot/orphan class. - acked-rename read-your-writes: a confirmed-materialized title is never silently lost/repainted. The verbs draw only WORLD entropy, never `self.rng` (the pick stream), so the reweight is the sole seed shift; calm replays byte-identically and the pinned chaos seeds converge + replay deterministically (re-verified, kept as-is per the #786 precedent). Swarm: chaos 0..200 + calm 0..100 x1500, faithful default — 200/200 + 100/100 green (both re-runs). Verification floor green: fmt, clippy -D warnings, nextest engram-dst 27/27, engram-sim 31/31 (sim legs), engram-coordinator 383/383. The operator-drain verb is NOT in the profile menu — driven only by dedicated tests (tests/api_surface.rs the real gRPC admin_drain_host; workload_verbs.rs the sequential cordon+evict). Folding it is blocked on two findings reported in the PR: (1) the evict pipeline's real-fs blob writes race the paused clock (the known in-memory-blob-store determinism issue), and (2) it uncovers a REAL capacity-soft evac over-reservation — evac_resumer's pick_for_session binds a measured-full survivor (the #722/#795 class on the dormant #775 evac leg), whose full fix needs reserved evac placement. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR
nikhilunni
added a commit
that referenced
this pull request
Jul 19, 2026
* R4: fold the API workload verbs into the swarm + two new model-oracle assertions ADR 0098 Phase 3 wave 4 — the deferred workload-verb fold-in from #786/#793. Folds the deterministic API verbs #786 wired (but drove only in dedicated tests) into BOTH swarm profiles at small weights, each over the real service surface and each feeding the acked-only expected-state model oracle: - Prompt: the full send_prompt (gRPC) -> real Deliver op -> run_started ack loop. Active-only + the ack closes the loop so an unacked outbox row cannot redeliver forever (no real guest harness emits the ack in the sim). - Rename: the coordinator-owned `set_session_suggested_title` write (the exact call the harness-title sink performs), store-direct — the real-wire harness route's event-bus publish races the paused clock under time advances. - Destroy: the real Destroy op + teardown, driven through the op pipeline like ResumeSession (the client-facing delete_session wait-loop can't step under the single-step-per-pick sim model; its full handler stays in api_surface). The model oracle grows two assertions (non-vacuity proven in tests/workload_verbs.rs): - acked-destroy-never-resurrects: a torn-down session that reappears live is the double-boot/orphan class. - acked-rename read-your-writes: a confirmed-materialized title is never silently lost/repainted. The verbs draw only WORLD entropy, never `self.rng` (the pick stream), so the reweight is the sole seed shift; calm replays byte-identically and the pinned chaos seeds converge + replay deterministically (re-verified, kept as-is per the #786 precedent). Swarm: chaos 0..200 + calm 0..100 x1500, faithful default — 200/200 + 100/100 green (both re-runs). Verification floor green: fmt, clippy -D warnings, nextest engram-dst 27/27, engram-sim 31/31 (sim legs), engram-coordinator 383/383. The operator-drain verb is NOT in the profile menu — driven only by dedicated tests (tests/api_surface.rs the real gRPC admin_drain_host; workload_verbs.rs the sequential cordon+evict). Folding it is blocked on two findings reported in the PR: (1) the evict pipeline's real-fs blob writes race the paused clock (the known in-memory-blob-store determinism issue), and (2) it uncovers a REAL capacity-soft evac over-reservation — evac_resumer's pick_for_session binds a measured-full survivor (the #722/#795 class on the dormant #775 evac leg), whose full fix needs reserved evac placement. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR * R4: per-seed hermetic runtime for the DST swarm — fix the cross-seed paused-clock CI hang The `simulation swarm (engram-dst)` lane hung (no progress, killed at the 6 h job timeout; normal ~5 min) after #798 was rebased onto main-with-#799. Root cause (harness, not product). The swarm binary ran EVERY seed of a `--seeds` range on ONE shared paused-clock current-thread runtime. The workload verbs drive the REAL coordinator handlers, which legitimately spawn DETACHED tokio tasks — the op executor's completion re-drive (`session_ops::enqueue`), the 15 s within-step op-heartbeat INTERVAL, the outbox Deliver op the prompt path fires. Some park on tokio timers, and `SimClock` IS tokio's paused clock (`advance` == `tokio::time::advance`), so a timer-parked detached task is a live participant in the sim's time model that the step loop's `drain_detached` (bare `yield_now`s, which never move virtual time) cannot reap. On the shared runtime those tasks LEAKED across the seed boundary and accumulated; a later seed's runtime then parked the OS thread (futex_wait — a deadlock, not a spin) waiting on the accumulated cross-seed timer/waker state instead of auto-advancing virtual time. Load/timing dependent: reproduced on the Linux dev VM under CPU contention (parked at seed ~126/155), intermittently on macOS under repeated full-window runs; NEVER when a seed runs in isolation (why per-seed local runs were green). Same class as #799 one level up — #799 removed the tokio::fs blob I/O; this removes the cross-seed runtime sharing that bridged seeds. Fix: - `engram_dst::run_seed` runs each seed on its OWN current-thread paused-clock runtime and DROPS it before the next — reaping every detached task at the boundary, so no state crosses seeds. This restores the seed independence a DST swarm requires, and is exactly what a single `--seed N` replay already does. The swarm binary and the multi-seed `first_sim` tests (same latent pattern) both route through it. - A per-seed wall-clock watchdog (own OS thread, 600 s budget) converts any residual/future liveness hang into a FAST, NAMED failure — a multi-hour CI hang is the worst reporting mode, so the swarm fails in minutes with the offending seed + its replay line, never at a 6 h job timeout. Verification: faithful chaos 0..200 -> 200/200, calm 0..100 -> 100/100 and byte-identical replay; the OLD binary deadlocks (parks forever) under load at the seed the FIXED binary sails past; macOS 60-run stress zero hangs (old binary hung); `nextest -p engram-dst -p engram-sim` 60/60 (incl. the pinned 722_r3 / 790 / 96 / 762 seeds and the reworked first_sim batches); fmt + clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR * R4: fix HostRegistry list/unbind_session holding a DashMap guard across .await (the seed-126 DST deadlock) The `simulation swarm (engram-dst)` lane WATCHDOG-timed-out at seed 126 (exactly 600 s of ZERO progress) on the 4-vcpu blacksmith CI runner — DESPITE the per-seed hermetic-runtime fix. A seed diverging by ENVIRONMENT after that fix means a within-seed real-world coupling remained. (The per-seed watchdog did its job: fast, named failure instead of a 6 h hang — which is how this was even findable.) Root cause (a real product concurrency bug, exposed by the single-threaded DST executor). `HostClient for HostRegistry::list` and `::unbind_session` fanned out over the fleet with: for entry in self.hosts.iter() { entry.value().backend.<rpc>().await; } DashMap's `iter()` holds the CURRENT entry's shard `RwLock` (read) across the whole loop body — INCLUDING the `.await`. On the sim's SINGLE-THREADED tokio current_thread runtime, while that iteration future is suspended at the await, the HostHeartbeats step's `HostRegistry::register` (a shard WRITE lock via `DashMap::insert`, fired every drain round) blocks the ONLY thread on that same shard — and the suspended iterator can never be polled to release its guard. Permanent deadlock. gdb of a hung process: the runtime thread parked in `dashmap … lock_exclusive_slow` under `HostRegistry::register`, 0 % CPU, `futex_wait`; the watchdog thread parked alongside. Which shard a `HostId` lands on is DashMap's per-process `RandomState` hash (getrandom-seeded) sized by `available_parallelism()` — so the collision (two HostIds sharing one shard) was a ~1.3 % getrandom-dependent, host-count- and CPU-count-dependent hang: measured 4/300 on the 16-core VM, far higher pinned to 1 cpu (4 shards), and the 4-vcpu blacksmith runner (16 shards) reliably rolled a colliding hash. This is "instance 3" of the determinism-audit class where a real-world OS coupling leaks into the paused-clock sim: #799 removed the `tokio::fs` blob I/O; the per-seed-runtime fix removed cross-seed runtime sharing; this removes the getrandom-seeded DashMap shard-lock deadlock. Fix (root, in-poll — the MemBlobStorage-analog: do the OS-coupled work synchronously first, hold nothing across the suspension). Snapshot the backends into a `Vec<Arc<dyn HostClient>>` FIRST — the `.collect()` fully drains the `iter()` and drops every shard guard — THEN `.await` per host over the owned Vec. This is the exact rule `backend_of`'s own doc already states ("cloning the trait object is cheap so callers don't have to hold the registry's internal entry across await points"); `list`/`unbind_session` were the two that violated it. It is a latent bug in production too — multi-threaded prod only STALLS a worker rather than permanently deadlocking, but holding a lock across an `.await` is always wrong. Verification: seed 126 x400 fresh processes (each a fresh getrandom hash; was 4/300 ≈ 1.3 % before) → ZERO hangs; seed 126 x150 pinned to 1 cpu / 4 shards (where the pre-fix binary hung heavily) → ZERO hangs; faithful chaos 0..200 → 200/200, calm 0..100 → 100/100; `nextest -p engram-coordinator -p engram-dst -p engram-sim` green; fmt + clippy clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Wave 2 Track A of the resiliency program — the three R2 rows of ADR 0098 §"Phase 3: the honest-gap program". Three logical commits over
crates/engram-dst(+engram-sim, + one contained coordinator determinism fix).Per-commit summary
1. The effect queue — in-flight interruption + message faults. Pre-R2 every
SimHostClientverb mutated world truth inline, so RPC loss/reorder/duplication and a replica crash BETWEEN the store commit (the ack the coordinator already holds) and the host-side world effect were structurally impossible (audit finding #2/#4). A mutating verb now records its world mutation as a typedEffect. Inline delivery is the default (byte-for-byte the pre-R2 world → Calm seeds unchanged, since Calm never opens a deferred window). When a host is in the deferred set (a Chaos fault window) the effect is queued under a monotonic serial (BTreeMap→ deterministic) for a later scheduler step:DeferHost/DeliverEffects(serial order) /DropEffect(loss) /DuplicateEffect/ReorderEffects(seeded shuffle). Crash/restart of a host severs its in-flight effects; quiescence closes every window and flushes in order before the fleet heals.tests/effect_queue.rspins the two now-reachable states (a deferred create withheld-then-delivered; a replica crash inside the commit→effect window followed by loss) both converging.2. API-driven workload over the real Router/gRPC surface.
workload.rsdrives each replica's ACTUAL surface — the tonicAppSessionService/AppFleetServiceimpls and the actual axumapi::routerviatower::oneshot— andtests/api_surface.rsexercises it deterministically (create acked + persisted + replays byte-identically; prompt+delete round-trip; the axum router runs the real HTTP handlers). SimMeta gains four PG-faithful methods the real handlers reach (were panic-stubs): broker-token insert/get/delete, teleport-target get/set,rebind_session_guarded,list_active_assignments_with_budgets_on_host, with broker-token + teleport conformance scenarios (ADR 0098 D4).3. The expected-state model oracle (the auditor). TigerBeetle's auditor shape: a model fed ONLY by acked outcomes, diffed against world/SimMeta truth in the standing invariant pass (every step + quiescence). Asserts acked-create/resume-never-silently-lost (the #570 class) and read-your-acked-writes (image never repainted). Non-vacuity proven in
tests/model_oracle.rs: drive a session to Active → drop its row directly → auditor FIRES → restore → clears.real-wire vs handler-direct
tower::oneshoton the realapi::routerruns the bearer middleware + extractorsharness_idleemitter (#775 trigger's wire)AppSessionServicetonic impl (auth + convert.rs + the same*_core) with a constructedtonic::Requestcarrying the bearer; only the socket/HTTP2 hop is skipped (a live tonic server would break the paused-clock, current-thread determinism the sim rests on)AppFleetService— the operator drainAuth reuses the test shape: the app-gRPC
BearerAuthis configured with a token every request carries (real auth check runs); the axum internal surface uses the empty-list dev bypass exactly liketests/api.rs::build_app. Neither path bypasses validation inside a handler.Model-boundary honesty note
The auditor is fed ONLY by acked outcomes (a session observed to reach
Activewith a durable row). Because of that, a lost-response op is legitimately absent from the model and is never asserted on — a create whose boot never durably established, or an op whose reply dropped, simply isn't a fact the model holds. This is the deliberate TigerBeetle boundary: the model records what the client was TOLD, not what the system attempted.Seed-shift notes
issue_722@600,seed_96@1500,issue_762@5000) were re-verified to still converge and are kept as-is; thewedged_bootpin is hand-driven (not seed-random) and is unaffected. Calm's pick table is untouched by commits 1 & 3, so Calm seeds are unchanged.Swarm results
Both CI windows, release build, zero failures:
--seeds 0..200 --steps 1500-> 200/200 OK--seeds 0..100 --steps 1500-> 100/100 OKThe effect-queue faults and the model auditor both hold across the full windows. fmt / clippy
-D warnings/ nextest on engram-dst + engram-sim + engram-coordinator (383/383) are green; hakari verifies.Findings (real bugs)
Determinism leak — FIXED here (contained).
create_session's prepare path (api/sessions.rs) minted the session id via a rawSessionId::new()(Uuid::new_v4) instead of the injectedservices.entropy— an ADR 0098 D1 violation; the API-create session id diverged every replay. In prodservices.entropyisOsEntropy(identical randomness), so behavior-identical there; the fix purely enables deterministic replay.Wire-skew / staging suppression + a driver-double-boot class — FOLLOW-UP (out of scope). Folding the API workload into the swarm requires the sim hosts to be genuinely schedulable (report the current
WIRE_VERSION, advertise the seeded image inready_images) and to produce recoverable snapshots. The pre-R2 sim left hosts wire-skewed and snapshots non-recoverable, which SILENTLY SUPPRESSED the digest-gated placement path the drivers use. Making the hosts faithful unmasks a real class the swarm fires on within a few hundred steps:single-ownership— a session ends up owning 2 live sandboxes across the fleet (ADR 0090 split-brain). Reproduces from the op-path workload ALONE once hosts are schedulable (a driver re-boots a session a create/resume already booted); confirmed independent of the API verbs.placement-accounting— Sigma reserved > allocatable on a host.quiescence-no-stragglers— an operator-drained session stuck atEvacuating.snapshot-safety— a session reachesIdlewithout a recoverable durable copy under evict->resume cycling.These are genuine coordinator (or deep sim-fidelity) findings needing their own investigation; shipping them would red the swarm lane, so this PR drives the real surface deterministically in dedicated tests instead. Recommended follow-up: land the host-fidelity fixes, root-cause the double-boot (suspected: the queue-scanner / OpReclaim orphan-backstop re-booting an already-booted session once candidates are non-empty), then fold the lifecycle-mutating API verbs +
#775harness-idle eviction + operator-drain into the swarm.Followups
🤖 Generated with Claude Code
https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR