Skip to content

R2: the #777 design calls — honest-Dead stage-2 predicate + ask-the-host bound-row policy - #782

Merged
nikhilunni merged 3 commits into
mainfrom
resilience/wave2-hostlost-policy
Jul 18, 2026
Merged

nikhilunni merged 3 commits into
mainfrom
resilience/wave2-hostlost-policy

Conversation

@nikhilunni

@nikhilunni nikhilunni commented Jul 18, 2026 •

Copy link
Copy Markdown
Contributor

Implements the two design calls recorded in #777 (final comment, user decision 2026-07-18), building on the #770 straggler sweep. One PR, two logical commits.

Commit 1 — honest-Dead: unify the HostLost stage-2 recoverability predicate

The two dead-host stage-2 sites (dead_host::evict_host and the #770 host_lost_straggler_sweep) keyed on mere snapshot presence (snapshot.is_some()), while reconcile::flip_missing already keyed on the recoverable flag. A session whose latest snapshot was un-recoverable (its BlobStorage HEAD failed at capture time) settled to a lying Idle a /resume could never honor. Per #777 the recoverable-filtered predicate is now THE stage-2 decision everywhere.

  • Shape: reused what flip_missing already calls — latest_snapshot_for_session(...), reading its .recoverable field. No MetadataStore method added or changed (least-churn honest shape). dead_host::recovery_target becomes the ONE shared pub(crate) predicate (keyed on has_recoverable_snapshot OR a live disk manifest), now called by all three stage-2 sites. That also closes flip_missing's reverse lie — it ignored the live disk manifest, so a disk-recoverable session could be routed to Dead.
  • Bad-capture signal: the "snapshot rows exist but none recoverable ⇒ Dead" arm gets a distinct warn! + new counter HOST_LOST_UNRECOVERABLE_SNAPSHOT_TOTAL (mirrors R0.3a: fix HostLost straggler (nightly seed 33043259, #762) #770's HOST_LOST_STRAGGLERS_SETTLED_TOTAL) at every stage-2 site.
  • D4 conformance: new latest_snapshot_reports_recoverable_flag scenario pins, across BOTH stores, that latest_snapshot_for_session returns the newest row and its recoverable flag faithfully — the store foundation the predicate rests on.
  • Red-then-green: engram-dst::host_lost_policy::unrecoverable_only_straggler_settles_dead_not_idle — a bound HostLost straggler whose only snapshot is recoverable=false settles Dead. Verified red first: with the pre-HostLost stage-2 recoverability predicate divergence: dead_host settles on ANY snapshot row, reconcile filters recoverable #777 snapshot.is_some() predicate it settled Idle (left: Idle, right: Dead), green after.

Commit 2 — bound rows: ask-the-host before destroying

The #770 sweep best-effort destroyed a still-bound sandbox after the 60s min-age. But a HostLost row can carry a live VM — a >60s heartbeat partition / binding desync parks the session at HostLost while the VMM is still serving (#776-review concern). Destroying it kills a live VM and rewinds resume to the last checkpoint.

Verification

  • cargo fmt --check + cargo clippy --all-targets -- -D warnings on engram-coordinator / engram-sim / engram-dst: clean.
  • cargo nextest run for those three crates: 421 passed, incl. both new red-then-green proofs and the pinned nightly sim: quiescence-no-stragglers violated #762 straggler seed (chaos 5000) still green.
  • Conformance across BOTH stores against the reachable dev Postgres (localhost:5435): t_latest_snapshot_reports_recoverable_flag::{sim,pg} + the neighbouring host-lost / snapshot scenarios all pass.
  • Swarm sanity 0..100 chaos × 1500: all 100 seeds converged, 0 regressions.

Refs #777, #770.

nikhilunni and others added 3 commits July 18, 2026 12:10
Before this, the two dead-host stage-2 sites keyed on mere snapshot
presence (`snapshot.is_some()`) while `reconcile::flip_missing` already
keyed on the `recoverable` flag — a session whose latest snapshot was
un-recoverable (its BlobStorage HEAD failed at capture time) settled to a
lying `Idle` that a `/resume` could never honor. Per the #777 design call
the recoverable-filtered predicate is now THE stage-2 decision everywhere.

- `dead_host::recovery_target` becomes the ONE shared predicate
  (`pub(crate)`), keyed on `has_recoverable_snapshot` (the flag, not row
  presence) OR a live disk manifest. `dead_host::evict_host`, the
  `host_lost_straggler_sweep`, AND `reconcile::flip_missing` all call it —
  which also closes flip_missing's reverse lie (it ignored the live disk
  manifest, so a disk-recoverable session could be routed to Dead).
- The "snapshot rows exist but none recoverable ⇒ Dead" arm gets a
  distinct warn + a new counter (`HOST_LOST_UNRECOVERABLE_SNAPSHOT_TOTAL`,
  mirroring the #770 stragglers-settled counter) at every stage-2 site, so
  a bad-capture pipeline stays visible instead of hiding behind Dead.
- D4 conformance: a new `latest_snapshot_reports_recoverable_flag`
  scenario pins, across BOTH stores, that `latest_snapshot_for_session`
  returns the newest row and its `recoverable` flag faithfully (the
  foundation the predicate rests on). No MetadataStore method changed.
- engram-dst proof `unrecoverable_only_straggler_settles_dead_not_idle`:
  a bound HostLost straggler whose only snapshot is `recoverable=false`
  now settles Dead. Proven red-then-green — with the pre-#777
  `snapshot.is_some()` predicate it settled Idle (assert: left=Idle,
  right=Dead), green after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The #770 straggler sweep best-effort DESTROYED a still-bound sandbox after
a 60s min-age. But a HostLost row can carry a LIVE VM — a >60s heartbeat
partition / binding desync parks the session at HostLost while the VMM is
still up and serving (the #776-review concern). Destroying it there kills
a live VM and rewinds any resume to the last checkpoint.

Per the #777 design call the sweep now consults HOST TRUTH before
destroying. The running-sandbox SET is not persisted in PG (only a count +
last_heartbeat_at), so — following the task's fallback — it uses the same
direct probe reconcile's ADR 0068 belt already uses: `probe_sandbox` →
`process_alive`, via the existing HostClient surface (no new RPC).

- If the host still reports the sandbox `process_alive`, the sweep DEFERS
  the destroy+settle for the reattach machinery and banks a per-session
  serving-strike (`HOST_LOST_STRAGGLER_DEFERRED_SERVING_TOTAL`). Only after
  `straggler_serving_strike_cap` (default 3, ~3 cycles) consecutive serving
  cycles — or a probe that fails / a not-alive process / no backend in the
  registry (host already gone) — does it destroy+settle as before.
- Convergence (oracle #8's shape) is preserved: every arm terminates in a
  settle within bounded cycles; the serving-defer removes the
  destroy-a-live-VM window without reintroducing the #762/#769 eternal
  wedge. Documented in the sweep's doc comment.
- Strikes are a per-replica in-memory `StragglerStrikeMap` threaded through
  `run_once` (mirroring `ProbeMemory`): no natural PG column exists, it's a
  per-pod backstop, and a settle by ANY replica ends the deferral for all
  via the #211 CAS + Conflict-idempotent transition. Pruned to the live
  HostLost set each cycle. engram-dst threads + clears it like probe_memory.
- engram-dst proof `serving_straggler_is_deferred_then_settles_at_the_strike_cap`:
  a bound HostLost straggler with a live world-side sandbox survives the
  first two sweeps (row HostLost, VM alive) and settles Idle on the third
  (cap reached, VM destroyed). Proven red-then-green — with the pre-#777
  best-effort destroy the first cycle killed the VM and settled (assert:
  left=Idle, right=HostLost), green after.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…cipline)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@nikhilunni

Copy link
Copy Markdown
Contributor Author

Orchestrator review: rebased onto current main (clean), verified 424/424 + fmt + clippy post-rebase, and swapped StragglerStrikeMap to BTreeMap (its only iteration is the prune, so no live leak — but BTreeMap is the discipline in decision-adjacent paths). The commit-1 reverse-lie fix (flip_missing ignoring live disk manifests) is a genuine catch beyond the #777 scope. Merging on green.

🤖 Generated with Claude Code

@nikhilunni
nikhilunni force-pushed the resilience/wave2-hostlost-policy branch from f859f9c to a9c48f0 Compare July 18, 2026 19:12
@nikhilunni
nikhilunni marked this pull request as ready for review July 18, 2026 19:26
@nikhilunni
nikhilunni merged commit 683b417 into main Jul 18, 2026
21 checks passed
@nikhilunni
nikhilunni deleted the resilience/wave2-hostlost-policy branch July 18, 2026 19:26
nikhilunni added a commit that referenced this pull request Jul 19, 2026
`run_evict_pipeline`'s composed path recorded the capture's
`recoverable` flag (a live `verify_snapshot_recoverable` BlobStorage
HEAD) and then flipped the session `Idle` unconditionally. A TRANSIENT
blob-HEAD blip stamps `recoverable = false` on a capture whose artifacts
are actually durable, so the session lands a lying `Idle` — the falsehood
surfaces only at `/resume` (which then fails → `Dead` via the #782 honest
predicate) after the user already tried to return.

Gate the terminal `Idle` transition on the same honest predicate the
dead-host stage-2 ladder uses (#782 `recovery_target`): a manifest-bearing
capture with `recoverable = false` and no live disk manifest is not a safe
basis for `Idle`. Abort the still-in-flight snapshot and return a RETRYABLE
error — no new state, no new RPC. The op machinery redrives within the
existing evict budget; a transient blip clears on the fresh capture+verify
→ `Idle` recoverable. A persistently-unrecoverable capture exhausts the
budget and falls back (nominated) to `HostLost`, where the dead-host
straggler sweep settles it `Dead` WITH the distinct unrecoverable-snapshot
signal against the recorded `recoverable = false` row — never Idle.

Scope is exactly the transient-HEAD-blip class (`manifests_present &&
!recoverable`). Manifest-less captures (Process/VZ dev backends; the
pre-#791 sim fidelity gap) keep their prior behavior — #791 already closed
the manifest-less case at the sim-fidelity layer, and gating it here would
wedge those backends into eternal retry. Only the idle path is gated.

Tests (idle_evictor): transient-then-success (HEAD fails once → guard
returns retryable, session stays Evicting + VM alive → redrive lands Idle
recoverable), persistent-failure (HEAD always fails → budget exhausts →
HostLost, never Idle-unrecoverable), and a scope test (manifest-less
capture still lands Idle). Backed by `engram_testkit::FaultyBlobStorage`
HEAD scripting + a manifest-bearing test backend.

No `MetadataStore`/PostgresStore SQL change → no D4 conformance delta. The
coordinator sim (engram-dst) has no storage-fault model, so a deterministic
transient blob-HEAD blip isn't expressible there yet (R3-storage-lies
territory); covered by the unit tests.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 19, 2026
`run_evict_pipeline`'s composed path recorded the capture's
`recoverable` flag (a live `verify_snapshot_recoverable` BlobStorage
HEAD) and then flipped the session `Idle` unconditionally. A TRANSIENT
blob-HEAD blip stamps `recoverable = false` on a capture whose artifacts
are actually durable, so the session lands a lying `Idle` — the falsehood
surfaces only at `/resume` (which then fails → `Dead` via the #782 honest
predicate) after the user already tried to return.

Gate the terminal `Idle` transition on the same honest predicate the
dead-host stage-2 ladder uses (#782 `recovery_target`): a manifest-bearing
capture with `recoverable = false` and no live disk manifest is not a safe
basis for `Idle`. Abort the still-in-flight snapshot and return a RETRYABLE
error — no new state, no new RPC. The op machinery redrives within the
existing evict budget; a transient blip clears on the fresh capture+verify
→ `Idle` recoverable. A persistently-unrecoverable capture exhausts the
budget and falls back (nominated) to `HostLost`, where the dead-host
straggler sweep settles it `Dead` WITH the distinct unrecoverable-snapshot
signal against the recorded `recoverable = false` row — never Idle.

Scope is exactly the transient-HEAD-blip class (`manifests_present &&
!recoverable`). Manifest-less captures (Process/VZ dev backends; the
pre-#791 sim fidelity gap) keep their prior behavior — #791 already closed
the manifest-less case at the sim-fidelity layer, and gating it here would
wedge those backends into eternal retry. Only the idle path is gated.

Tests (idle_evictor): transient-then-success (HEAD fails once → guard
returns retryable, session stays Evicting + VM alive → redrive lands Idle
recoverable), persistent-failure (HEAD always fails → budget exhausts →
HostLost, never Idle-unrecoverable), and a scope test (manifest-less
capture still lands Idle). Backed by `engram_testkit::FaultyBlobStorage`
HEAD scripting + a manifest-bearing test backend.

No `MetadataStore`/PostgresStore SQL change → no D4 conformance delta. The
coordinator sim (engram-dst) has no storage-fault model, so a deterministic
transient blob-HEAD blip isn't expressible there yet (R3-storage-lies
territory); covered by the unit tests.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 20, 2026
…bridge-fault exploration (#784) (#807)

* R7: the shared NBD device-plane model + the coordinator rehydrate-list core (#784)

ADR 0098 R-CoSim rung 2 foundation. Extract the P7/#806 device-plane world
model — generation, the real NbdSlotAllocator, served_by/kernel_owner/parked,
the guest_holds_device proof-of-death input — into a pub, reusable
`engram_dst_host::device_plane::DevicePlane` keyed by SandboxId. Every decision
delegates to the REAL host-core verdicts (sweep_verdict incl. the #806 holder
table, resume_data_plane_served, is_local_survivor_candidate); every slot lease
comes from the REAL NbdSlotAllocator. The disk backend stays the owning host's
concern (rebuilt after `serve` reports a newly-served device) so the plane is
pure device bookkeeping.

Also extract `register_rehydrate_list_core` from the coordinator's `register`
handler (the run_once-for-handlers pattern rung 1 established) so the
co-simulated host's register-rehydrate leg drives the EXACT listing the real
coordinator names.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* R7: co-simulate the NBD device-plane family at the boundary + capture-lock pin (#784)

ADR 0098 R-CoSim rung 2, deliverables 1 + 5. Port the device plane into the
CosimHost (via the shared DevicePlane, keyed by the coordinator-minted
SandboxIds — rung 1's documented divergence), and drive the full survivor
family co-simulated, all real code on both sides:

- host-agent roll (new generation; survivors resident, kernel_owner persists
  at the dead gen);
- register-rehydrate against the REAL coordinator listing
  (register_rehydrate_list_core over the bridge): coord-list pass → #739 local
  ChainHeadRecord pass → stale-binding sweep;
- the REAL sweep_verdict incl. the #806 (liveness × holder) table — a
  dead-owner device a live guest still holds is PARKed, never severed;
- the un-pause data-plane gate (resume_data_plane_served);
- the coordinator's REAL host_lost_straggler_sweep with the #782 probe/strike
  arms — including the #777 tension end-to-end across the boundary (a bound
  HostLost row whose VM probes ALIVE past the 60s min-age is DEFERRED via
  ask-the-host, then settles at the strike cap; a departed VM settles at once).

Directed pins: the fixed happy path (listed survivor re-served, un-pause
lands), the #806 ungated PARK-then-reserve-zero-loss, the un-pause gate firing
on a disconnected dead-guest plane, and both straggler-sweep arms. Standing
oracles added: severed-live-holder (#806), slot-accounting (P7 #3), and
cross-boundary ownership agreement.

Capture-lock release pin (deliverable 5): a QUARANTINED eviction finalize
releases the capture lock (capture_in_flight → false), driven through the REAL
run_eviction_finalize_attempt by sabotaging the staging dir — so #783's
teardown-reconcile capture-in-flight exemption can never become a permanent
reap-shield.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* R7: the seeded boundary swarm + bridge faults + standing oracles (#784)

ADR 0098 R-CoSim rung 2, deliverables 2 + 3. A `--seeds` swarm over the
coordinator↔host boundary (`sim-cosim` binary + the in-lane `tests/swarm.rs`
replay-twice/convergence checks), small world (1 replica, 1 host, ≤3 sessions
— the state-space product is the risk; bounded and documented). The scheduler
interleaves REAL coordinator drivers, REAL host steps, host-agent rolls, and
bridge faults (the applied-commit-but-ack-lost window reached via explicit
perturbations that mirror the effect outcomes: DropHostBinding = lost bind ack,
ForceHostLost = the partition window, ForceTerminal = lost destroy).

Determinism (audit items 6-8): per-seed runtime dropped at the boundary + the
named 600s watchdog (the engram-dst run_seed shape); world entropy only;
BTreeMap/index-ordered picks; coarse explicit time steps so no fs-I/O clock
auto-advance can tip a decision threshold. Replay-twice is byte-identical
(verified in-lane + via the CLI).

Standing oracles checked each step + at quiescence: severed-live-holder (#806),
slot-accounting no-plane-leak (P7 #3), the ACTIVE-serve cross-boundary
ownership split-brain guard, idle⇒durable-snapshot (#570, at quiescence
post-finalize-drain), teardown/ownership completeness, and bounded convergence
(every session terminal-or-stable, no Evicting/HostLost wedge).

Three swarm firings RCA'd honestly while landing (each a real cross-system
finding or a fidelity gap, no oracle weakened):
- ownership oracle fired on a terminal session's served device (teardown
  window, not a re-home split-brain) and on an evicted sandbox mid-finalize
  (paused, benign): scoped the every-step guard to ACTIVELY-serving planes
  (live backend, non-terminal), added the STRONG quiescence completeness
  closure so the terminal/eviction cases stay covered;
- #570 idle-durability fired on the D5-Idle-before-async-finalize window (a
  real prod window): moved to quiescence post-finalize-drain (detection of a
  cancelled-finalize loss preserved);
- Evicting-wedge: a Roll decoupled from rehydrate left an unrehydrated Active
  session (unreachable in prod — a roll's startup always rehydrates); coupled
  Roll → register-rehydrate faithfully.
- next_free_device spare-lease overlap (a real device-double-allocation bug in
  the shared plane) fixed.

PR window: 0..40 x 600 chaos+calm, ~1.2s (per-seed ~29ms), well under 5 min.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* R7: wire the boundary swarm into CI + the nightly cosim-swarm job (#784)

ADR 0098 R-CoSim rung 2, deliverable 4. The existing `test-cosim` lane gains a
fixed PR swarm window (0..40 x 600 chaos + calm, ~1.2s/window, well under 5
min); nightly-sim.yml gains a `cosim-swarm` job with date-derived
non-overlapping windows (200 chaos + 80 calm x 1500, sized below the host
swarm's 480/night since the co-sim runs both sides per step) whose
--failure-report feeds the same slug-deduped issue filing with a "nightly
cosim:" title prefix. CI Gate already needs test-cosim (unchanged). YAML
validated with python3 yaml.safe_load.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* R7: ADR 0098 bookend — R-CoSim rung 2 (the boundary device-plane family + swarm) (#784)

Record the Wave 7 arc: the reusable DevicePlane extraction (+ the deferred
SimHost migration, honestly noted), the co-simulated roll→rehydrate→sweep→
un-pause + straggler #777 family, the seeded boundary swarm with its
determinism/I/O-audit verdict and small-world bounding, and the four swarm
firings each RCA'd without weakening an oracle.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* R7: in-memory blob store for the co-sim (determinism, #784)

Swap both blob tiers to the deterministic in-memory engram_sim::MemBlobStorage
(audit items 6/7): the host's ChunkStore (was a per-run tempdir LocalBlobStorage)
and — more importantly — the coordinator replica's blob/chunk_store, which shared
a PROCESS-GLOBAL temp_dir()/engram-dst-cosim-blobs across every SimWorld (a
cross-world contamination + real-fs-latency clock-drift hazard, the #793/#799
class, latent under the directed tests but fatal to a multi-seed swarm). The
ChunkedDiskBackend's own ChunkCache still uses the SimFs tempdir (inherent to
reusing the REAL backend); its I/O never feeds a decision (the swarm's coarse
explicit time crosses every threshold). Replay-twice stays byte-identical.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant