Skip to content

R2 Track A: effect queue + Router-driven workload + the expected-state model oracle - #786

Merged
nikhilunni merged 3 commits into
mainfrom
resilience/wave2-track-a
Jul 18, 2026
Merged

nikhilunni merged 3 commits into
mainfrom
resilience/wave2-track-a

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

Wave 2 Track A of the resiliency program — the three R2 rows of ADR 0098 §"Phase 3: the honest-gap program". Three logical commits over crates/engram-dst (+ engram-sim, + one contained coordinator determinism fix).

Per-commit summary

1. The effect queue — in-flight interruption + message faults. Pre-R2 every SimHostClient verb mutated world truth inline, so RPC loss/reorder/duplication and a replica crash BETWEEN the store commit (the ack the coordinator already holds) and the host-side world effect were structurally impossible (audit finding #2/#4). A mutating verb now records its world mutation as a typed Effect. Inline delivery is the default (byte-for-byte the pre-R2 world → Calm seeds unchanged, since Calm never opens a deferred window). When a host is in the deferred set (a Chaos fault window) the effect is queued under a monotonic serial (BTreeMap → deterministic) for a later scheduler step: DeferHost / DeliverEffects (serial order) / DropEffect (loss) / DuplicateEffect / ReorderEffects (seeded shuffle). Crash/restart of a host severs its in-flight effects; quiescence closes every window and flushes in order before the fleet heals. tests/effect_queue.rs pins the two now-reachable states (a deferred create withheld-then-delivered; a replica crash inside the commit→effect window followed by loss) both converging.

2. API-driven workload over the real Router/gRPC surface. workload.rs drives each replica's ACTUAL surface — the tonic AppSessionService/AppFleetService impls and the actual axum api::router via tower::oneshot — and tests/api_surface.rs exercises it deterministically (create acked + persisted + replays byte-identically; prompt+delete round-trip; the axum router runs the real HTTP handlers). SimMeta gains four PG-faithful methods the real handlers reach (were panic-stubs): broker-token insert/get/delete, teleport-target get/set, rebind_session_guarded, list_active_assignments_with_budgets_on_host, with broker-token + teleport conformance scenarios (ADR 0098 D4).

3. The expected-state model oracle (the auditor). TigerBeetle's auditor shape: a model fed ONLY by acked outcomes, diffed against world/SimMeta truth in the standing invariant pass (every step + quiescence). Asserts acked-create/resume-never-silently-lost (the #570 class) and read-your-acked-writes (image never repainted). Non-vacuity proven in tests/model_oracle.rs: drive a session to Active → drop its row directly → auditor FIRES → restore → clears.

real-wire vs handler-direct

verb path why
admin pause / resume real-wire HTTP a route exists; tower::oneshot on the real api::router runs the bearer middleware + extractors
harness-event ingest real-wire HTTP host->coord ingestion route; the harness_idle emitter (#775 trigger's wire)
create_session / send_prompt / resume / delete_session handler-direct gRPC ADR 0051 retired the REST create — gRPC-only, no axum route to oneshot. Driven through the real AppSessionService tonic impl (auth + convert.rs + the same *_core) with a constructed tonic::Request carrying the bearer; only the socket/HTTP2 hop is skipped (a live tonic server would break the paused-clock, current-thread determinism the sim rests on)
admin_drain_host handler-direct gRPC AppFleetService — the operator drain

Auth reuses the test shape: the app-gRPC BearerAuth is configured with a token every request carries (real auth check runs); the axum internal surface uses the empty-list dev bypass exactly like tests/api.rs::build_app. Neither path bypasses validation inside a handler.

Model-boundary honesty note

The auditor is fed ONLY by acked outcomes (a session observed to reach Active with a durable row). Because of that, a lost-response op is legitimately absent from the model and is never asserted on — a create whose boot never durably established, or an op whose reply dropped, simply isn't a fact the model holds. This is the deliberate TigerBeetle boundary: the model records what the client was TOLD, not what the system attempted.

Seed-shift notes

  • Commit 1 re-weights the Chaos pick table (carves 9 points for the effect-queue arms). This shifts every Chaos seed's exploration (seeds pin to a commit, per the D6 precedent). The three pinned chaos seeds (issue_722 @600, seed_96 @1500, issue_762 @5000) were re-verified to still converge and are kept as-is; the wedged_boot pin is hand-driven (not seed-random) and is unaffected. Calm's pick table is untouched by commits 1 & 3, so Calm seeds are unchanged.
  • Commits 2 & 3 do NOT touch the swarm pick tables (see the scope note in Findings), so no further seed shift.

Swarm results

Both CI windows, release build, zero failures:

  • chaos --seeds 0..200 --steps 1500 -> 200/200 OK
  • calm --seeds 0..100 --steps 1500 -> 100/100 OK

The effect-queue faults and the model auditor both hold across the full windows. fmt / clippy -D warnings / nextest on engram-dst + engram-sim + engram-coordinator (383/383) are green; hakari verifies.

Findings (real bugs)

  1. Determinism leak — FIXED here (contained). create_session's prepare path (api/sessions.rs) minted the session id via a raw SessionId::new() (Uuid::new_v4) instead of the injected services.entropy — an ADR 0098 D1 violation; the API-create session id diverged every replay. In prod services.entropy is OsEntropy (identical randomness), so behavior-identical there; the fix purely enables deterministic replay.

  2. Wire-skew / staging suppression + a driver-double-boot class — FOLLOW-UP (out of scope). Folding the API workload into the swarm requires the sim hosts to be genuinely schedulable (report the current WIRE_VERSION, advertise the seeded image in ready_images) and to produce recoverable snapshots. The pre-R2 sim left hosts wire-skewed and snapshots non-recoverable, which SILENTLY SUPPRESSED the digest-gated placement path the drivers use. Making the hosts faithful unmasks a real class the swarm fires on within a few hundred steps:

    • single-ownership — a session ends up owning 2 live sandboxes across the fleet (ADR 0090 split-brain). Reproduces from the op-path workload ALONE once hosts are schedulable (a driver re-boots a session a create/resume already booted); confirmed independent of the API verbs.
    • placement-accounting — Sigma reserved > allocatable on a host.
    • quiescence-no-stragglers — an operator-drained session stuck at Evacuating.
    • eviction snapshot-safety — a session reaches Idle without a recoverable durable copy under evict->resume cycling.
      These are genuine coordinator (or deep sim-fidelity) findings needing their own investigation; shipping them would red the swarm lane, so this PR drives the real surface deterministically in dedicated tests instead. Recommended follow-up: land the host-fidelity fixes, root-cause the double-boot (suspected: the queue-scanner / OpReclaim orphan-backstop re-booting an already-booted session once candidates are non-empty), then fold the lifecycle-mutating API verbs + #775 harness-idle eviction + operator-drain into the swarm.

Followups

  • The double-boot / wire-skew-suppression follow-up above (highest value — a latent ADR 0090 violation the fidelity fixes reveal).
  • Extend the auditor with acked-op-result-consistency + completed-run-repaint checks once the API workload is swarm-safe (they need acked op replies the op-path feeding doesn't surface).

🤖 Generated with Claude Code

https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR

nikhilunni and others added 3 commits July 18, 2026 12:57
…faults)

Pre-R2 every SimHostClient verb mutated world truth inline after one
optional delay, so RPC loss/reorder/duplication and a replica crash
BETWEEN the store commit (the ack the coordinator already holds) and the
host-side world effect were structurally impossible (the audit's finding
#2/#4).

A mutating verb now records its world mutation as a typed `Effect`.
Inline delivery is the default — byte-for-byte the pre-R2 world, so Calm
seeds are unchanged (Calm never opens a deferred window). When the target
host is in the deferred set (a Chaos fault window), the effect is queued
under a monotonic serial (BTreeMap → deterministic order) for a later
scheduler step to deliver / drop / duplicate / reorder:

- DeferHost(i, on)   — open/close a host's deferred window
- DeliverEffects     — deliver the queue in serial (causal) order
- DropEffect         — loss: drop one queued effect (seeded pick)
- DuplicateEffect    — re-deliver one effect (seeded pick)
- ReorderEffects     — deliver in a seeded-shuffled order

Crash/restart of a host severs its in-flight effects (never resurrected
onto the cleared VM set); quiescence closes every window and flushes the
queue in order before the fleet heals, so world truth is consistent with
the coordinator's commits. The Chaos weight table carves 9 points out of
AdvanceTime/Driver/HostHeartbeats/Crash/RestartHost for the new arms,
which shifts every Chaos seed's exploration (seeds pin to a commit); the
pinned chaos seeds still converge and are kept.

tests/effect_queue.rs pins the two reachable states the queue unlocks: a
deferred create withheld-then-delivered, and a replica crash inside the
commit→effect window followed by effect loss, both converging at
quiescence (no_op_dropped + no-stragglers + no-orphans).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Pre-R2 the workload called store/core fns directly (0/17 HTTP, 0/~69 gRPC
driven — the audit's finding). This lands the real surface:

- workload.rs drives each replica's ACTUAL surface: the tonic
  AppSessionService/AppFleetService impls (auth + convert.rs + the same
  *_core the axum handler calls — handler-direct, no socket, to keep the
  paused-clock current-thread determinism) for create/prompt/resume/
  delete/drain, and the actual axum api::router via tower::oneshot
  (real-wire, middleware + extractors) for admin pause/resume and the
  harness-idle host ingestion. A module honesty table documents which
  verb takes which path.
- tests/api_surface.rs exercises it deterministically: create is acked,
  persisted (never lost), and REPLAYS byte-identically; prompt+delete
  round-trip; the axum router runs the real HTTP handlers to a normal
  response.

Real bug found + fixed: create_session's prepare path minted the session
id via a raw `SessionId::new()` (Uuid::new_v4) instead of the injected
`services.entropy` — an ADR 0098 D1 determinism leak (the id diverged
every replay; in prod OsEntropy makes this behavior-identical). The
API-create path now replays.

SimMeta gains four PG-faithful methods the real handlers reach
(previously panic-stubs): insert/get/delete_broker_token,
get/set_teleport_target, rebind_session_guarded,
list_active_assignments_with_budgets_on_host — with broker-token and
teleport-target conformance scenarios (ADR 0098 D4).

Scope note: folding this workload into the chaos/calm SWARM additionally
needs host-fidelity fixes (schedulable wire_version, staged ready_images)
that unmask a genuine, separate driver-double-boot + eviction-durability
class the pre-R2 sim silently suppressed by leaving hosts wire-skewed and
snapshots non-recoverable. Those are real findings for a follow-up; here
the surface is driven directly and deterministically rather than shipping
a red swarm lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
TigerBeetle's auditor shape: an expected-state model fed ONLY by ACKED
workload outcomes, diffed against world/SimMeta truth in the standing
invariant pass (every step + at quiescence). The honesty boundary is
explicit — because the model is fed only by acks, a lost-response op (a
create whose boot never durably established, an op whose reply dropped)
is legitimately absent from the model and never asserted on.

What it asserts:
- acked-create/resume never silently lost: a session the workload saw
  reach Active (durable row committed) still has a row on every later
  step, unless a later ACKED destroy retired it. A row vanishing under a
  live session is exactly the #570 symptom class (coordinator-unbind vs
  host-teardown-reconcile racing a session out of existence) and the
  durability-lie class this program exists to catch.
- read-your-acked-writes: the image a session was created with is never
  repainted (both replicas read the same shared SimMeta, so a present row
  is readable on either).

Fed from the op-path CreateSession/ResumeSession outcomes (record_if_live
records a session once observed Active). Wired into run()'s per-step and
quiescence checks alongside the existing oracles.

Non-vacuity is proven, not assumed: tests/model_oracle.rs drives a
session to Active, drops its row directly (the corruption the #570 race
produces), and asserts the auditor FIRES with `model-acked-create-not-
lost`; restoring the row clears it. `SimMetadataStore::with_db_mut` is the
test-only corruption injector.

Swarm (chaos 0..40 x600) stays green — the auditor is a safety invariant
the current code upholds; the canary test is what proves it can bite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@nikhilunni

Copy link
Copy Markdown
Contributor Author

Orchestrator review: the sessions.rs entropy fix is behavior-identical in prod (OsEntropy) and closes a live D1 leak in a driven path — the exact hole the R1.7b ids.rs proposal predicted. The scope call (dedicated deterministic API tests now; swarm integration held until #787's double-boot is fixed) is the right honesty tradeoff — a red lane teaches nothing. The double-boot finding is filed as #787 with your reproduction notes and marked as the next wave's opener. Merging on green.

🤖 Generated with Claude Code

@nikhilunni
nikhilunni marked this pull request as ready for review July 18, 2026 21:29
@nikhilunni
nikhilunni merged commit 66887ec into main Jul 18, 2026
22 checks passed
@nikhilunni
nikhilunni deleted the resilience/wave2-track-a branch July 18, 2026 21:29
nikhilunni added a commit that referenced this pull request Jul 19, 2026
… assertions

ADR 0098 Phase 3 wave 4 — the deferred workload-verb fold-in from #786/#793.

Folds the deterministic API verbs #786 wired (but drove only in dedicated
tests) into BOTH swarm profiles at small weights, each over the real service
surface and each feeding the acked-only expected-state model oracle:

- Prompt: the full send_prompt (gRPC) -> real Deliver op -> run_started ack
  loop. Active-only + the ack closes the loop so an unacked outbox row cannot
  redeliver forever (no real guest harness emits the ack in the sim).
- Rename: the coordinator-owned `set_session_suggested_title` write (the exact
  call the harness-title sink performs), store-direct — the real-wire harness
  route's event-bus publish races the paused clock under time advances.
- Destroy: the real Destroy op + teardown, driven through the op pipeline like
  ResumeSession (the client-facing delete_session wait-loop can't step under
  the single-step-per-pick sim model; its full handler stays in api_surface).

The model oracle grows two assertions (non-vacuity proven in
tests/workload_verbs.rs):
- acked-destroy-never-resurrects: a torn-down session that reappears live is
  the double-boot/orphan class.
- acked-rename read-your-writes: a confirmed-materialized title is never
  silently lost/repainted.

The verbs draw only WORLD entropy, never `self.rng` (the pick stream), so the
reweight is the sole seed shift; calm replays byte-identically and the pinned
chaos seeds converge + replay deterministically (re-verified, kept as-is per
the #786 precedent).

Swarm: chaos 0..200 + calm 0..100 x1500, faithful default — 200/200 + 100/100
green (both re-runs). Verification floor green: fmt, clippy -D warnings,
nextest engram-dst 27/27, engram-sim 31/31 (sim legs), engram-coordinator
383/383.

The operator-drain verb is NOT in the profile menu — driven only by dedicated
tests (tests/api_surface.rs the real gRPC admin_drain_host; workload_verbs.rs
the sequential cordon+evict). Folding it is blocked on two findings reported
in the PR: (1) the evict pipeline's real-fs blob writes race the paused clock
(the known in-memory-blob-store determinism issue), and (2) it uncovers a REAL
capacity-soft evac over-reservation — evac_resumer's pick_for_session binds a
measured-full survivor (the #722/#795 class on the dormant #775 evac leg),
whose full fix needs reserved evac placement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR
nikhilunni added a commit that referenced this pull request Jul 19, 2026
* R4: fold the API workload verbs into the swarm + two new model-oracle assertions

ADR 0098 Phase 3 wave 4 — the deferred workload-verb fold-in from #786/#793.

Folds the deterministic API verbs #786 wired (but drove only in dedicated
tests) into BOTH swarm profiles at small weights, each over the real service
surface and each feeding the acked-only expected-state model oracle:

- Prompt: the full send_prompt (gRPC) -> real Deliver op -> run_started ack
  loop. Active-only + the ack closes the loop so an unacked outbox row cannot
  redeliver forever (no real guest harness emits the ack in the sim).
- Rename: the coordinator-owned `set_session_suggested_title` write (the exact
  call the harness-title sink performs), store-direct — the real-wire harness
  route's event-bus publish races the paused clock under time advances.
- Destroy: the real Destroy op + teardown, driven through the op pipeline like
  ResumeSession (the client-facing delete_session wait-loop can't step under
  the single-step-per-pick sim model; its full handler stays in api_surface).

The model oracle grows two assertions (non-vacuity proven in
tests/workload_verbs.rs):
- acked-destroy-never-resurrects: a torn-down session that reappears live is
  the double-boot/orphan class.
- acked-rename read-your-writes: a confirmed-materialized title is never
  silently lost/repainted.

The verbs draw only WORLD entropy, never `self.rng` (the pick stream), so the
reweight is the sole seed shift; calm replays byte-identically and the pinned
chaos seeds converge + replay deterministically (re-verified, kept as-is per
the #786 precedent).

Swarm: chaos 0..200 + calm 0..100 x1500, faithful default — 200/200 + 100/100
green (both re-runs). Verification floor green: fmt, clippy -D warnings,
nextest engram-dst 27/27, engram-sim 31/31 (sim legs), engram-coordinator
383/383.

The operator-drain verb is NOT in the profile menu — driven only by dedicated
tests (tests/api_surface.rs the real gRPC admin_drain_host; workload_verbs.rs
the sequential cordon+evict). Folding it is blocked on two findings reported
in the PR: (1) the evict pipeline's real-fs blob writes race the paused clock
(the known in-memory-blob-store determinism issue), and (2) it uncovers a REAL
capacity-soft evac over-reservation — evac_resumer's pick_for_session binds a
measured-full survivor (the #722/#795 class on the dormant #775 evac leg),
whose full fix needs reserved evac placement.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR

* R4: per-seed hermetic runtime for the DST swarm — fix the cross-seed paused-clock CI hang

The `simulation swarm (engram-dst)` lane hung (no progress, killed at the 6 h
job timeout; normal ~5 min) after #798 was rebased onto main-with-#799.

Root cause (harness, not product). The swarm binary ran EVERY seed of a
`--seeds` range on ONE shared paused-clock current-thread runtime. The workload
verbs drive the REAL coordinator handlers, which legitimately spawn DETACHED
tokio tasks — the op executor's completion re-drive (`session_ops::enqueue`),
the 15 s within-step op-heartbeat INTERVAL, the outbox Deliver op the prompt
path fires. Some park on tokio timers, and `SimClock` IS tokio's paused clock
(`advance` == `tokio::time::advance`), so a timer-parked detached task is a live
participant in the sim's time model that the step loop's `drain_detached` (bare
`yield_now`s, which never move virtual time) cannot reap. On the shared runtime
those tasks LEAKED across the seed boundary and accumulated; a later seed's
runtime then parked the OS thread (futex_wait — a deadlock, not a spin) waiting
on the accumulated cross-seed timer/waker state instead of auto-advancing
virtual time. Load/timing dependent: reproduced on the Linux dev VM under CPU
contention (parked at seed ~126/155), intermittently on macOS under repeated
full-window runs; NEVER when a seed runs in isolation (why per-seed local runs
were green). Same class as #799 one level up — #799 removed the tokio::fs blob
I/O; this removes the cross-seed runtime sharing that bridged seeds.

Fix:
- `engram_dst::run_seed` runs each seed on its OWN current-thread paused-clock
  runtime and DROPS it before the next — reaping every detached task at the
  boundary, so no state crosses seeds. This restores the seed independence a DST
  swarm requires, and is exactly what a single `--seed N` replay already does.
  The swarm binary and the multi-seed `first_sim` tests (same latent pattern)
  both route through it.
- A per-seed wall-clock watchdog (own OS thread, 600 s budget) converts any
  residual/future liveness hang into a FAST, NAMED failure — a multi-hour CI
  hang is the worst reporting mode, so the swarm fails in minutes with the
  offending seed + its replay line, never at a 6 h job timeout.

Verification: faithful chaos 0..200 -> 200/200, calm 0..100 -> 100/100 and
byte-identical replay; the OLD binary deadlocks (parks forever) under load at
the seed the FIXED binary sails past; macOS 60-run stress zero hangs (old binary
hung); `nextest -p engram-dst -p engram-sim` 60/60 (incl. the pinned 722_r3 /
790 / 96 / 762 seeds and the reworked first_sim batches); fmt + clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR

* R4: fix HostRegistry list/unbind_session holding a DashMap guard across .await (the seed-126 DST deadlock)

The `simulation swarm (engram-dst)` lane WATCHDOG-timed-out at seed 126 (exactly
600 s of ZERO progress) on the 4-vcpu blacksmith CI runner — DESPITE the per-seed
hermetic-runtime fix. A seed diverging by ENVIRONMENT after that fix means a
within-seed real-world coupling remained. (The per-seed watchdog did its job:
fast, named failure instead of a 6 h hang — which is how this was even findable.)

Root cause (a real product concurrency bug, exposed by the single-threaded DST
executor). `HostClient for HostRegistry::list` and `::unbind_session` fanned out
over the fleet with:

    for entry in self.hosts.iter() { entry.value().backend.<rpc>().await; }

DashMap's `iter()` holds the CURRENT entry's shard `RwLock` (read) across the
whole loop body — INCLUDING the `.await`. On the sim's SINGLE-THREADED tokio
current_thread runtime, while that iteration future is suspended at the await,
the HostHeartbeats step's `HostRegistry::register` (a shard WRITE lock via
`DashMap::insert`, fired every drain round) blocks the ONLY thread on that same
shard — and the suspended iterator can never be polled to release its guard.
Permanent deadlock. gdb of a hung process: the runtime thread parked in
`dashmap … lock_exclusive_slow` under `HostRegistry::register`, 0 % CPU,
`futex_wait`; the watchdog thread parked alongside.

Which shard a `HostId` lands on is DashMap's per-process `RandomState` hash
(getrandom-seeded) sized by `available_parallelism()` — so the collision (two
HostIds sharing one shard) was a ~1.3 % getrandom-dependent, host-count- and
CPU-count-dependent hang: measured 4/300 on the 16-core VM, far higher pinned to
1 cpu (4 shards), and the 4-vcpu blacksmith runner (16 shards) reliably rolled a
colliding hash. This is "instance 3" of the determinism-audit class where a
real-world OS coupling leaks into the paused-clock sim: #799 removed the
`tokio::fs` blob I/O; the per-seed-runtime fix removed cross-seed runtime
sharing; this removes the getrandom-seeded DashMap shard-lock deadlock.

Fix (root, in-poll — the MemBlobStorage-analog: do the OS-coupled work
synchronously first, hold nothing across the suspension). Snapshot the backends
into a `Vec<Arc<dyn HostClient>>` FIRST — the `.collect()` fully drains the
`iter()` and drops every shard guard — THEN `.await` per host over the owned
Vec. This is the exact rule `backend_of`'s own doc already states ("cloning the
trait object is cheap so callers don't have to hold the registry's internal
entry across await points"); `list`/`unbind_session` were the two that violated
it. It is a latent bug in production too — multi-threaded prod only STALLS a
worker rather than permanently deadlocking, but holding a lock across an `.await`
is always wrong.

Verification: seed 126 x400 fresh processes (each a fresh getrandom hash; was
4/300 ≈ 1.3 % before) → ZERO hangs; seed 126 x150 pinned to 1 cpu / 4 shards
(where the pre-fix binary hung heavily) → ZERO hangs; faithful chaos 0..200 →
200/200, calm 0..100 → 100/100; `nextest -p engram-coordinator -p engram-dst
-p engram-sim` green; fmt + clippy clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PBLK8qJSNJQK722E1n2omR

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant