Skip to content

chore(web)(deps): bump framer-motion from 11.18.2 to 12.38.0 in /web - #4

Merged
nikhilunni merged 1 commit into
mainfrom
dependabot/npm_and_yarn/web/framer-motion-12.38.0
May 10, 2026
Merged

nikhilunni merged 1 commit into
mainfrom
dependabot/npm_and_yarn/web/framer-motion-12.38.0

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github May 10, 2026 •

Copy link
Copy Markdown
Contributor

Bumps framer-motion from 11.18.2 to 12.38.0.

Changelog

Sourced from framer-motion's changelog.

[12.38.0] 2026-03-16

Added

  • Added layoutAnchor prop to configure custom anchor point for resolving relative projection boxes.

Fixed

  • Reorder: Fix axis switching after window resize.
  • Reorder: Fix with virtualised lists.
  • AnimatePresence: Ensure children are removed when exit animation matches current values.

[12.37.0] 2026-03-16

Added

  • Support for hardware accelerating "start" and "end" offsets in scroll and useScroll.
  • Support for oklch, oklab, lab, lch, color, color-mix, light-dark color types.

Fixed

  • Fix whileInView with client-side navigation.
  • Fix draggable elements when layout updates due to surrounding element re-renders.
  • Improved memory pressure of layout animations.
  • Ensure motion value returned from useSpring reports correct isAnimating().

[12.36.0] 2026-03-09

Added

  • Allow dragSnapToOrigin to accept "x" or "y" for per-axis snapping.
  • Added axis-locked layout animations with layout="x" and layout="y".
  • Added skipInitialAnimation to useSpring.

Fixed

  • Fixed height and width: auto animations with box-sizing: border-box.
  • Reset component values when exit animation finishes.
  • Ensure anticipate easing returns 1 at p === 1.
  • Fix @emotion/is-prop-valid resolve error in Storybook.
  • Remove data-pop-layout-id from exiting elements when animation interrupted.
  • Ensure we skip WAAPI for non-animatable keyframes.
  • Ensure we skip WAAPI for SVG transforms.
  • Ensure MotionValue props are not passed to SVG.
  • AnimatePresence: Prevent mode="wait" elements from getting stuck when switched rapidly.

[12.35.2] 2026-03-09

Fixed

... (truncated)

Commits
  • 0bfc9fe v12.38.0
  • 343cb0c Updating layoutAnchor
  • ee99ad2 Updating changelog
  • 062660b Updating changgelog
  • 303da7d Updating readme
  • b075adc Merge pull request #3647 from motiondivision/feat/layout-anchor
  • f0991d6 Add missing layoutAnchor !== false guard in attemptToResolveRelativeTarget
  • b5798e9 Merge pull request #3642 from motiondivision/worktree-fix-issue-3078
  • 7686c19 Merge pull request #3636 from motiondivision/worktree-fix-issue-3061
  • a95c487 Fix auto-scroll in reorder-virtualized test page
  • Additional commits viewable in compare view

@dependabot dependabot Bot added dependencies Pull requests that update a dependency file javascript Pull requests that update javascript code labels May 10, 2026
@nikhilunni

Copy link
Copy Markdown
Contributor

@dependabot recreate

@dependabot
dependabot Bot force-pushed the dependabot/npm_and_yarn/web/framer-motion-12.38.0 branch from a82fe27 to 9c021f4 Compare May 10, 2026 22:39
Bumps [framer-motion](https://github.com/motiondivision/motion) from 11.18.2 to 12.38.0.
- [Changelog](https://github.com/motiondivision/motion/blob/main/CHANGELOG.md)
- [Commits](motiondivision/motion@v11.18.2...v12.38.0)

---
updated-dependencies:
- dependency-name: framer-motion
  dependency-version: 12.38.0
  dependency-type: direct:production
  update-type: version-update:semver-major
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot
dependabot Bot force-pushed the dependabot/npm_and_yarn/web/framer-motion-12.38.0 branch from 9c021f4 to 5eaff14 Compare May 10, 2026 22:45
@nikhilunni
nikhilunni merged commit 1442242 into main May 10, 2026
9 checks passed
@nikhilunni
nikhilunni deleted the dependabot/npm_and_yarn/web/framer-motion-12.38.0 branch May 10, 2026 22:51
nikhilunni added a commit that referenced this pull request May 12, 2026
Tier 4 #3(a). PooledBackend gains `with_chunk_cache(cache)`; when
present, `materialize_chunked_rootfs` routes chunk reads through
the cache via `materialize_to_file_cached`. Chunks shared across
manifests (canonical-base images, fork lineage) now serve from
local NVMe on subsequent materializes instead of round-tripping
BlobStorage every time.

Wired in both binaries:
- Coordinator `--mode=all`: cache root at
  `<local_path>/chunk-cache/`, default 200 GiB budget.
- Standalone host-agent: same default at `<work_dir>/chunk-cache/`.

HostAgent grows `with_chunk_cache(cache)` that forwards into
PooledBackend during `run()`. A new
`materialize_chunked_rootfs_uses_chunk_cache_when_present` test
proves the cache path is exercised: after warming, the test
deletes the underlying chunks from the blob store and asserts the
second materialize still succeeds — it can only succeed if the
chunks were served from the cache.

Materialized-file orphan reap (3(b)) and chunk-store GC scheduler
(#4) still pending — both pair naturally and need the same
admin-endpoint + cron pattern.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request May 12, 2026
Tier 4 #4 — explicit-trigger chunk-store GC. Pairs with the
project's "implicit-trigger features need a paired admin
endpoint firing the same primitive" pattern (see disk-pressure /
flush). Cron scheduler lands later; the admin endpoint is the
testable + drain-before-redeploy ops shape.

Surface:
- POST /api/admin/gc-chunks?retain_secs=N (default 86400 = 24h)
  triggers a single GC pass. Returns GcChunksResult JSON with
  chunks_deleted, bytes_freed, chunks_retained_age, elapsed_ms,
  live_manifest_count, retain_secs.
- New MetadataStore::list_live_disk_manifest_ids() returns
  DISTINCT manifest_ids referenced by any snapshot row. Backed by
  the idx_snapshots_disk_manifest partial index from migration
  0018 — fast scan even on busy deployments.
- Services.chunk_store: Arc-cheap clone of the same store the
  --mode=all PooledBackend already builds. Single source of truth
  per process.

What the GC live-set DOES NOT yet cover:
- enabled_images' canonical manifests (no snapshot row points
  at them until someone takes a snapshot using that image).
  Workaround: operators pass a generous retain_secs window so
  newly-baked images survive the GC gap. Real fix: persist
  disk_manifest_ref on enabled_images and union the two sets.
  Tracked in docs/chunked-storage-rollout.md.

Verification:
- 586/586 workspace nextest pass.
- Two unit tests in tests/api.rs lock the empty-store response
  shape + the retain_secs query-param round-trip.
- New live-PG test admin_gc_chunks_live_pg seeds a snapshot row
  with a manifest, fires the endpoint, asserts live_manifest_count
  reflects the seeded ref. CI runs it via the Postgres-gated
  step.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request May 12, 2026
Phase 5 follow-up #32 slice 1: plumb the per-image bake-time
canonical memory manifest through the type system end-to-end.
The image-builder doesn't yet auto-populate it (slice 2 of #32
implements the bake-time FC boot+pause+chunk dance), but the
wiring is in place — once an operator pre-computes a canonical
ref on a sidecar Linux+KVM job, the field flows cleanly from
`bundle.json` → `CachedImage.bundle.canonical_memory_manifest`
→ `SandboxSpec.canonical_memory_manifest` → `FcSnapshotManifest.
canonical_memory_manifest` → UFFD handler's `--canonical-manifest`.

Plumbing:
- `engram_core::types::sandbox::SandboxSpec.canonical_memory_manifest`
  (new, `#[serde(default)]` — pre-existing specs round-trip).
- `engram_host_agent::image_cache::ImageBundle.canonical_memory_manifest`
  (new, serde-default).
- `engram_image_builder::BuildRequest.canonical_memory_manifest` —
  optional pre-computed ref; written to `bundle.json` verbatim.
- `engram_sandbox_firecracker::FcSnapshotManifest.
  canonical_memory_manifest` — now sourced from
  `spec.canonical_memory_manifest` at snapshot time (was
  hardcoded `None`). UFFD restore sees the canonical ref from
  the FC sidecar JSON the way it always has.
- `PooledBackend::create()` lifts
  `cached.bundle.canonical_memory_manifest` onto
  `spec.canonical_memory_manifest` after image-cache resolve.

Mass-edit: 18 `SandboxSpec { ... }` literals across the workspace
(coord, sandbox-process, sandbox-vz, sandbox-firecracker tests,
host-agent, protocol) gain `canonical_memory_manifest: None`.
Same for 5 `BuildRequest` literals and 2 `ImageBundle` test
constructors.

ADR 0007 e2e test suite (`crates/engram-host-agent/tests/
adr_0007_e2e.rs`):
- **#1 — snapshot path patches sidecar JSON's memory_manifest +
  chunks memory.bin into the store**: fake FC backend writes the
  on-disk shape a real FC snapshot produces; PooledBackend.snapshot
  chunks the emitted memory.bin and patches the JSON. Asserts
  metadata.memory_manifest is set, the JSON sidecar carries the
  matching ref, and `materialize_to_file` round-trips byte-for-byte.
- **#2 — cross-host restore materialises memory.bin from chunks**:
  host A creates + snapshots; the local memory.bin is then deleted
  to simulate cross-host transfer (only state.bin + manifest.json
  staged); host B's PooledBackend (sharing the chunk store) calls
  restore — the wrap materialises memory.bin from the chunked
  memory_manifest before inner.restore sees it. Asserts the inner
  backend sees a fully-formed snapshot dir + bytes round-trip.
- **#3 — trace_host_hint round-trips through sidecar JSON**: the
  snapshotting host's id flows through serde so a cross-host
  restoring backend can pass `--prefault-trace <hint>` to the UFFD
  handler.
- **#4 — bundle.canonical_memory_manifest lifts onto SandboxSpec**:
  validates the wire shape so the canonical ref survives bundle
  → spec → snapshot manifest plumbing.
- **#5 — content-addressed dedup across sessions**: same bytes
  chunked twice produce identical hashes (the property the
  cross-session dedup story relies on).
- **#6 — Linux + nbd-module-gated NBD daemon spawn** (`#[ignore]`,
  validates the disk-side end-to-end on the dev VM): builds a
  disk manifest, attaches a real NBD daemon against `/dev/nbd0`,
  drops cleanly.

Pure Rust, no FC, no KVM, no real kernel — runs in 0.42s on
every CI lane (macOS + Linux). The Linux+KVM-gated tests for
real FC + real NBD live in `crates/engram-sandbox-firecracker/
tests/snapshot_uffd.rs` + the planned `nbd_chunked_disk.rs`
(task #34).

206 unit + integration tests pass workspace-wide:
- engram-host-agent lib: 85
- engram-host-agent adr_0007_e2e: 5
- engram-sandbox-firecracker lib: 38
- engram-image-builder lib: 15
- engram-core lib: 25
- engram-coordinator lib: 43
- (the new test file adds 5 + 1 ignored)

clippy + fmt clean on macOS; Linux check clean via dev VM.
nikhilunni added a commit that referenced this pull request May 12, 2026
- known-issues #9 (NBD) marked resolved with commit hashes
- known-issues #10 (UFFD-from-chunks) marked resolved with the
  shipped surface enumerated (canonical capture, working-set
  R&R, cross-host materialize-from-chunks)
- known-issues #13 (materialized-rootfs leak) marked resolved —
  reap_materialize_dir + admin endpoint + chunk_gc cron driver
- rollout doc Phase 1 GC scheduler ⬜ → ✅ + Tier 4 #4 ditto
- ADR 0007 "What this ADR does NOT cover" rewritten: NBD,
  UFFD, materialize-orphan-reap, 0018/0019 schema reshape all
  move from "not implemented" to shipped; observability + the
  Phase 6 destructive trait reshape remain the named gaps.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request May 20, 2026
Extends the Known Issues block with a "Pointers for the next
session" subsection so a fresh session can pick up without
re-deriving where to start. Per-issue:
- entry-point file paths + grep anchors
- proposed module additions (coord sweeper for #3, gauge for #4,
  heartbeat field for #5)
- a SQL snippet for verifying stale templates against BlobStorage

Plus:
- suggested attack order (#3 first to unblock warm-path test,
  then #1+#2 as one commit, then #5, then #4 as defense)
- ready-to-paste Cloud Logging queries for refill failures,
  snapshot dir leaks, eviction failures
- production state snapshot at write time so a future reader
  can diff against current state (FC hosts pgs9+tf4j on
  3b6aec3, demo image at warm-75babf7, harness still on
  pre-M1.12 b9dd2d1, 5+ known stale templates)
- pointers to engrams-prod-ops skill scripts

No code change — purely documentation hygiene for the handoff.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request May 20, 2026
ADR 0014 issue #4. Defense-in-depth backstop for the prod incident on
`engrams-fc-xngk` (99 GB disk filled in ~13 min at ~25 leaked snapshot
dirs × 4 GiB). Commits 1+2 (strict cascade + commit/abort lifecycle)
already prevent the specific leak class that triggered the original
incident, but if some future bug introduces a different on-disk leak,
we want the host's idle-evict driver to stop adding fuel to the fire
before disk fills.

Mechanism: on every idle-evict tick (default 10s), the host-side
driver in lib.rs calls `idle_evictor::disk_pressure_check(work_dir,
floor)`. If statvfs reports free bytes below the floor, the tick:
- increments `engram_host_idle_evict_disk_pressure_holds_total`,
- logs a WARN with free / floor for ops to page on,
- skips pushing candidates this tick.

`disk_pressure_check` fails open on statvfs error — we'd rather over-
evict than block all evictions silently on a transient FS hiccup.

Default floor: 20 GiB. Sized for ~5 concurrent in-flight idle-evicts
× ~4 GiB per FC memory dump + headroom (the actual `engrams-fc-xngk`
post-mortem shows that 4 GiB-per-attempt at the unbounded rate fills
the 99 GB disk in 13 min; 20 GiB free is the safety margin under
which we stop pushing new candidates). Tunable via
`ENGRAM_IDLE_EVICT_DISK_FLOOR_BYTES`.

New host-agent metrics:
- `engram_host_disk_free_bytes` (gauge) — sampled each tick.
- `engram_host_idle_evict_disk_pressure_holds_total` (counter) —
  increments each tick we hold off.

Tests added:
- `free_disk_bytes_returns_value_for_tempdir` + `_none_for_missing_path`
  cover the statvfs probe's Ok/Err branches.
- `disk_pressure_check_fails_open_on_statvfs_error` — missing path
  fails open (allow=true).
- `disk_pressure_check_allows_when_floor_is_zero` and
  `disk_pressure_check_blocks_when_floor_exceeds_capacity` cover
  both sides of the gate.

All 761 tests pass; `just check` clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 4, 2026
…(ADR 0037 P4d)

Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off
⇒ every restore takes today's cold start_agent path; a warm snapshot's
harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert
in prod).

- proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness
  (delivers session_id + session_env + first_prompt over the existing harness
  vsock channel). Threaded through the gRPC client/server, LocalHostClient
  (→ HarnessHub::bind), and HostRegistry routing.
- coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND
  the base snapshot's warm_harness. When set, instead of start_agent:
  apply_egress_policy + late_bind_harness(session_env merged with the harness
  extras — forge/upload tokens, owner — + first prompt). The warm claude is
  not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is
  what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_-
  PLACEHOLDER → real (reusing the token resolved for the cold per-session
  entry, swapping only the placeholder). warm_bind_enabled() kill-switch.

KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell
keep the capture-time (sentinel/placeholder) session env — FC merge_session_env
is a documented no-op for the guest env. The agent loop is correct (forge/
upload tokens reach the helpers via the P4b session-env file; OAuth via the
baked constant placeholder + proxy); a warm-path agentd env-merge RPC for
shell/exec attribution is a follow-up.

dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt
clean. End-to-end FC warm-capture→restore→bind validation is next (P5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 4, 2026
…ow-up

Bookend update: P4b (no-respawn bind + token file), P4c (warm capture),
P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e +
measurement) and P6 (Accepted) remain — they need a real-FC warm path run
(new hub-based e2e scaffolding) + a prod canary for the Accept-gate
latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge
gap (pitfall #4) as a follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 4, 2026
…(ADR 0037 P4d)

Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off
⇒ every restore takes today's cold start_agent path; a warm snapshot's
harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert
in prod).

- proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness
  (delivers session_id + session_env + first_prompt over the existing harness
  vsock channel). Threaded through the gRPC client/server, LocalHostClient
  (→ HarnessHub::bind), and HostRegistry routing.
- coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND
  the base snapshot's warm_harness. When set, instead of start_agent:
  apply_egress_policy + late_bind_harness(session_env merged with the harness
  extras — forge/upload tokens, owner — + first prompt). The warm claude is
  not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is
  what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_-
  PLACEHOLDER → real (reusing the token resolved for the cold per-session
  entry, swapping only the placeholder). warm_bind_enabled() kill-switch.

KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell
keep the capture-time (sentinel/placeholder) session env — FC merge_session_env
is a documented no-op for the guest env. The agent loop is correct (forge/
upload tokens reach the helpers via the P4b session-env file; OAuth via the
baked constant placeholder + proxy); a warm-path agentd env-merge RPC for
shell/exec attribution is a follow-up.

dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt
clean. End-to-end FC warm-capture→restore→bind validation is next (P5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 4, 2026
…ow-up

Bookend update: P4b (no-respawn bind + token file), P4c (warm capture),
P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e +
measurement) and P6 (Accepted) remain — they need a real-FC warm path run
(new hub-based e2e scaffolding) + a prod canary for the Accept-gate
latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge
gap (pitfall #4) as a follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 5, 2026
…(ADR 0037 P4d)

Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off
⇒ every restore takes today's cold start_agent path; a warm snapshot's
harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert
in prod).

- proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness
  (delivers session_id + session_env + first_prompt over the existing harness
  vsock channel). Threaded through the gRPC client/server, LocalHostClient
  (→ HarnessHub::bind), and HostRegistry routing.
- coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND
  the base snapshot's warm_harness. When set, instead of start_agent:
  apply_egress_policy + late_bind_harness(session_env merged with the harness
  extras — forge/upload tokens, owner — + first prompt). The warm claude is
  not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is
  what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_-
  PLACEHOLDER → real (reusing the token resolved for the cold per-session
  entry, swapping only the placeholder). warm_bind_enabled() kill-switch.

KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell
keep the capture-time (sentinel/placeholder) session env — FC merge_session_env
is a documented no-op for the guest env. The agent loop is correct (forge/
upload tokens reach the helpers via the P4b session-env file; OAuth via the
baked constant placeholder + proxy); a warm-path agentd env-merge RPC for
shell/exec attribution is a follow-up.

dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt
clean. End-to-end FC warm-capture→restore→bind validation is next (P5).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 5, 2026
…ow-up

Bookend update: P4b (no-respawn bind + token file), P4c (warm capture),
P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e +
measurement) and P6 (Accepted) remain — they need a real-FC warm path run
(new hub-based e2e scaffolding) + a prod canary for the Accept-gate
latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge
gap (pitfall #4) as a follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 8, 2026
Define the K-phase labels (used across chart comments + PRs but never
written down) and map them to the migration-path steps. Add design
sections for K3 (HostFleet CRD + drain-gated operator), K4 (demand
autoscaling off the coordinator capacity signal), and K5 (parallel-run
cutover + node-asset staging + MIG retirement). Record the stable-HostId
completion in K2; mark open questions 1-2 resolved by the K3/K4 designs
and refine #4 (pod-netns settled; dedicated tainted pool + PSA namespace
+ KVM device plugin as the remaining security posture).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jun 25, 2026
…e leak) (#439)

Each enabled-image re-bake/refresh captures a fresh per-image base
snapshot and swaps `enabled_images.base_snapshot_id` to it
(`upsert_enabled_image`'s ON CONFLICT), leaving the PRIOR base row
dangling: `session_id IS NULL`, referenced by no `enabled_images` row.
Nothing deleted it — `prune_session_snapshots` (checkpoint retention) is
`session_id IS NOT NULL` only — and an orphan base keeps pinning its own
disk+memory chunks via pin-set sources #3/#4
(`list_recoverable_snapshot_{disk,memory}_manifests`, which filter on
`recoverable = TRUE` with no session predicate). So every image refresh
permanently leaked one base snapshot's chunks (20-32 GB for the heavy
dogfood images). Prod had 23 such orphans (~150 GB logical) going back 9
days, none in any GC queue.

Add `prune_orphan_base_snapshots`, the `session_id IS NULL` mirror of
`prune_session_snapshots`: delete base rows older than
`ENGRAM_BASE_SNAPSHOT_RETENTION_HOURS` (default 24) that are referenced by
no `enabled_images.base_snapshot_id` (live OR soft-deleted — soft-deleted
lineage is intentionally still chunk-pinned, ADR 0021 P1.8), bumping
`chunk_generation` in the same TX (GC-barrier symmetry). The existing
chunk-GC (ADR 0016 Phase C) and snapshot-blob-GC (ADR 0028 addendum)
sweeps then reclaim the now-unpinned chunks and portable `snapshots/<id>/`
blobs — the reaper deletes nothing in BlobStorage directly. The
`base_snapshot_id` FK (REFERENCES snapshots(id), no ON DELETE) is a hard
backstop against ever deleting an in-use base.

Spawned beside `checkpoint_retention`. A sweeper (not an inline
delete-old-base in RefreshImage) so it both drains the existing backlog
and survives restarts, per the "scanner drives transitions" convention.

Also fixes the now-stale `snapshot_blob_pin_set` doc that claimed base
rows are never deleted.

Test (live-PG, wired into the existing Postgres-gated CI lane): a
superseded orphan past grace is reaped; the current (enabled-image) base
and a fresh orphan within grace are kept; session snapshots are
untouched. The sibling checkpoint-retention test's template is pinned
recent so the new global reaper can't collect it under local parallel
runs (CI serializes the lane).

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 3, 2026
…ss (finding #4)

Both are zero-caller writers left behind by reserve_and_persist_create
subsuming the old satellite-write paths: the sealed-secrets bytea and the
harness selection now ride the one-transaction write-set / row INSERT
directly (engram-postgres/src/lib.rs's reserve_and_persist_create), not
these standalone upserts. Confirmed zero callers workspace-wide (including
orchestrator/ and web/) before deleting — only the trait declarations, the
PostgresStore impls, and 10 mock impls referenced them.

Per the repo's clean-break convention, retire both from the trait, the PG
impl, and every mock rather than leaving them as an orphaned, unused
re-entry point for the FK-ordering/partial-write bug class this PR set out
to kill. get_session_secrets/delete_session_secrets and get_session_harness
are untouched — those remain live (read/delete) call sites.
nikhilunni added a commit that referenced this pull request Jul 3, 2026
…licated ~60-line copy)

PR #556 review finding #4 [CONFIRMED].

`api::prompt`'s emit-ordering tests carried a ~60-line copy of the exact
`AppState`/`Services`/`MiniMeta` wiring `api::snapshot`'s
`evicting_gate_tests` already had, justified by a comment claiming test
modules "can't cheaply share private test-only fns" — false for
same-crate unit tests, as `state::tests::MiniMeta` itself already
demonstrates being shared across files.

Hoists one `pub(crate) fn build_state_for_session` next to `MiniMeta` in
`state::tests`; both call sites now delegate to it (`api::snapshot`'s
copy trims to a 2-tuple wrapper since its tests don't need the `MiniMeta`
handle). Every future `Services` field addition now has one call site to
update, not two.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc
nikhilunni added a commit that referenced this pull request Jul 6, 2026
… overlapped boot legs + prompt-over-wire (#566)

* feat(coord): migration 0077 — session.selected_skills + fleet_catalog_changed NOTIFY trigger

Part of issue #535 (create-as-a-plan). Groundwork for two later commits:

- sessions.selected_skills (TEXT[]) persists a create's dynamic-mount
  selection so the queue scanner's boot re-prepare can reconstruct it
  (ADR 0055 TODO(P1-D): queued creates currently boot with base skills
  only, since the queue row never carried the selection).
- notify_fleet_catalog_changed() + a trigger on hosts scoped to
  current_bundles changes (guarded by IS DISTINCT FROM, so it stays
  quiet across the few-seconds heartbeat UPDATE and only fires on an
  actual host-roll stamp change) backs the coordinator's boot-bundle
  cache invalidation, added in a follow-up commit.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* feat(coord): per-enabled-image boot bundle cache (issue #535 a)

Combines the issue's steps 2+3 into one commit (the NOTIFY listener's
invalidation calls need the cache type to exist first; splitting them
would leave a dead intermediate compile state).

- New `boot_bundle` module: `BootBundleCache` read-through caches, per
  enabled image, the parsed manifest + fetched base-snapshot record +
  resolved memory/vcpu budgets (today re-derived on every create), plus
  the fleet's baked bundle-name→sha catalog (today re-scanned via
  `list_active_hosts` up to three times per create). Both are TTL'd
  (30s belt-and-braces against a dropped PgListener notification).
- `engram-postgres`: `upsert_enabled_image` / `soft_delete_enabled_image`
  now fire `pg_notify('enabled_image_changed', image_uri)` inside their
  existing transaction (delivered iff it commits); `delete_enabled_image`
  fires it best-effort after, mirroring `org_secret_changed`.
- `pg_listener`: subscribes `enabled_image_changed` (invalidates one
  cache entry) and `fleet_catalog_changed` (invalidates the whole
  catalog — see migration 0077's trigger, landed in the prior commit).
- `prepare_from_grpc` / `prepare_from_row` / `prepare_inner` /
  `fleet_bundle_catalog` rewired onto the cache: the strict
  (non-soft-deleted) vs. tolerant (`_any`) split moves to the two call
  sites (the cache always fills via the tolerant view), and
  `boot_on_reserved_host`'s per-create `get_snapshot` is gone — the
  snapshot record now rides `BootInputs.base_snapshot` from the bundle.

Net: a warm-cache create now does zero `toml::from_str` calls and zero
extra `list_active_hosts` scans beyond the one placement still needs
(`candidates_for` — deliberately kept, per the issue's "conscious
divergence": placement needs a heartbeat-fresh host view).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* feat(coord): one-transaction session write-set (issue #535 b)

Collapses `reserve_placement` + `enqueue_session_create` + the boot/
enqueue paths' separate satellite-write chains into a single
`MetadataStore::reserve_and_persist_create` call whose Postgres impl
commits the ENTIRE write-set — the row (placed or queued) plus every
satellite (sealed secrets, capabilities, integration policy, harness,
selected skills) — in ONE `FOR UPDATE` transaction, before any host RPC.

- `engram-core`: new `SessionCreateWriteSet` / `CreateDisposition` +
  `MetadataStore::reserve_and_persist_create` (replaces `reserve_
  placement` + `enqueue_session_create`) and `transition_session_created`
  (a slim `pending → created` + `sandbox_id` UPDATE, replacing `create_
  session_created`'s INSERT-or-UPDATE upsert — the row is now guaranteed
  to already exist). `Session` gains `selected_skills: Vec<String>`,
  fixing the ADR 0055 TODO(P1-D) gap: a queued create's boot re-prepare
  can now reconstruct its dynamic-mount selection instead of silently
  dropping to base skills.
- `engram-postgres`: the transactional impl (extends `reserve_placement`'s
  FOR-UPDATE body); `get_session`/`list_active_sessions`/
  `list_queued_sessions_fifo` project the new column.
- `engram-coordinator`: `boot_prepared` seals secrets (KEK, pure crypto —
  has no place inside the DB transaction) and serializes the policy
  BEFORE calling `reserve_and_persist_create`, then dispatches on
  `CreateDisposition` — `enqueue_create` as a separate function is gone,
  its Queued-disposition handling folds into `boot_prepared`.
  `boot_on_reserved_host` now does exactly ONE write of its own
  (`transition_session_created`, since the sandbox doesn't exist until
  the restore RPC returns) — the FK-ordering bug class (a satellite
  write racing the row's own insert; the ADR 0051 forge-token
  regression) is dead by construction, not by "call it after the row"
  convention.
- Every other `MetadataStore` impl (9 test/mock fixtures across 6
  crates) updated: the 2 that exercise the real create path (coordinator
  HTTP + gRPC integration tests) got honest in-memory equivalents; the
  rest mirror their pre-existing `unreachable!()`/`unimplemented!()`
  convention for unexercised trait surface.
- New Postgres-level tests (`placement_reservation_live_pg.rs`) proving
  the write-set's atomicity: a single call commits every satellite
  together, and a forced mid-transaction failure (duplicate session_id)
  leaves NOTHING from that attempt — not even satellites that would
  have followed the failing statement.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* feat(coord): overlap the independent boot legs with the restore RPC (issue #535 c)

`boot_on_reserved_host` no longer serializes the restore RPC behind the
env/egress work (or vice versa) — the two are independent (neither
touches the other's inputs) and now run concurrently via `tokio::join!`:

- Restore leg: `restore_base_on_host` (the ~0.4-0.7s VM-side work).
- Env/egress leg: the per-spawn forge/upload broker-token mint
  (`inject_harness_env`) + the integration policy's Plane-B injection
  resolution (`resolve_inject_entries`, which can round-trip an
  external mint-provider API for a mint-mode connector) — this is
  where that external round trip moves OFF the serial tail. Both only
  need `session_id`/`image_ref`/`integration_policy`, not the sandbox;
  the broker-token FK has been satisfiable since `reserve_and_persist_
  create` committed the row, well before this function runs.

`build_egress_policy` splits accordingly: `resolve_inject_entries` +
`build_observe_entries` (sandbox-independent, now called from the
overlapped leg) stay as-is; the renamed `assemble_egress_policy` is the
remaining sandbox-dependent half (`guest_ip` + final assembly).

Also parallelizes `resolve_policy_secrets`' per-secret `SecretStore`
round trips (order-insensitive — no secret depends on another) via
`futures::future::join_all`, replacing the one-at-a-time loop.

`transition_session_created` + the removal of `boot_on_reserved_host`'s
satellite writes already landed in the prior commit (they're the same
underlying change as the one-transaction write-set — splitting them
would have left a dead intermediate compile state), so this commit is
scoped to the actual leg-overlap + secret-resolution parallelization.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* feat(coord,harness): prompt over the wire (issue #535 d)

The initial prompt no longer rides `ENGRAM_INITIAL_PROMPT` env — every
prompt, first or follow-up, is now a harness-protocol `Prompt` frame:

- `api/prompt.rs`: factored the echo-then-forward core out of
  `send_prompt_core` into `deliver_prompt(state, session_id, sandbox_id,
  prompt_id, text)` — the user-echo-first ordering (load-bearing for web
  rendering) and the self-healing `deliver_with_reattach` forward, shared
  by every caller.
- `session_boot::boot_on_reserved_host`: after the Active flip, mints a
  server-side `prompt_id` and calls `deliver_prompt` for the initial
  prompt — replacing the synthetic `prompt_id: None` event. Delivery
  failure past the reattach budget is `BootError::Started` (terminal,
  consistent with a `start_agent` failure): a session that can't receive
  the prompt that created it is broken.
- `resolve_harness` no longer takes an `initial_prompt` param or inserts
  `ENGRAM_INITIAL_PROMPT`; `git grep ENGRAM_INITIAL_PROMPT` now returns
  nothing.
- `engram-harness-claude`: `run_engine` drops the `initial_prompt` param
  — the pending queue starts empty and the first prompt arrives via
  `HarnessCommand::Prompt` like every other one. Updated the 8 unit
  tests that seeded an initial prompt through the deleted parameter to
  instead send it via `cmd_tx` post-spawn (and to expect the leading
  `Idle` the engine now emits before any prompt arrives, since the
  env-seeded fast-path — "start turn 1 with no leading Idle" — no
  longer exists).
- `engram-host-agent/tests/e2e_harness.rs`: `capture_sink` now also
  returns a command sender so `drive_harness` can push the initial
  prompt as a wire frame instead of an env var — a hand-rolled minimal
  stand-in for `HarnessHub` (this test drives `SandboxBackend` directly,
  no coordinator/hub in the loop).

The queued path unifies for free: `SessionCreateWriteSet::queue_prompt`
(landed in the write-set commit) is already the durable prompt, and
`prepare_from_row` threads it into the identical `boot_on_reserved_host`
delivery path — no separate queued-prompt spelling.

Deviation: did not extend `e2e_stack.rs`'s create-with-prompt test
(`e2e_claude_with_bogus_key_surfaces_anthropic_auth_error`) to assert
`prompt_id` threads onto `run_started` — it's quarantined (#403,
excluded from the gating e2e lane) and requires a live
`ENGRAM_E2E_GRPC_ADDR` stack this environment doesn't have, so the
change is unverifiable here. The same property (RunStarted.prompt_id
matches the delivered prompt_id) is covered by engram-harness-claude's
`queue_holds_edits_and_consumes_type_ahead` unit test instead.

Note: engram-harness-claude's test module is `#[cfg(target_os =
"linux")]`-gated and e2e_harness.rs is Linux+KVM+FC+Docker+sudo-gated —
neither compiles or runs on this macOS dev machine; both are verified by
careful reading + will run for real in CI's Linux/FC lanes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* feat(coord): coord_prepare/coord_finalize phase metrics + doc pass (issue #535)

Adds the two phase labels the issue's acceptance criteria need to turn
"the coordinator serial tail is ~1s" from an estimate into a
measurement, on the existing `engram_session_boot_seconds` histogram
(no new metric names):

- `coord_prepare`: `create_session_core` entry through `reserve_and_
  persist_create`'s commit — the serial coordinator-side prefix ahead
  of the (now-concurrent, host-side) restore work. Recorded on the
  Placed path only.
- `coord_finalize`: the restore RPC returning through the `created →
  active` transition — the coordinator-owned tail after the host hands
  back a live sandbox. Success path only.

`total` minus (`coord_prepare` + `coord_finalize`) is the actual
host-side restore RPC wall time — the split this issue's evidence
section was missing.

Doc pass: `session_boot.rs`'s module header now describes the (a)-(d)
pipeline shape instead of the pre-refactor procedure; the FK-ordering
guard comment in `prepare_inner` (the anchor the issue tracked as
`sessions.rs:1197-1203`, drifted slightly by the time this landed)
rewritten to describe the current dead-by-construction invariant
instead of the historical hazard.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc

* fix(coord): guard transition_session_created on status=pending (finding #2)

The UPDATE that flips a session row to `created` + binds `sandbox_id` had
no status guard and ignored `rows_affected`, so it returned `Ok(())` even
when the row was deleted (DeleteSession) or requeued (stale-pending
scanner) while the restore RPC that precedes this call was in flight.
That silently binds a live sandbox onto a gone/inconsistent row instead
of hitting the existing `Err` teardown arm in session_boot.rs, which
destroys the now-orphaned sandbox.

Add `AND status = 'pending'` to the WHERE clause and return
MetaError::NotFound on rows_affected() == 0, matching the convention used
elsewhere in this file (e.g. delete_registry_credential).

* fix(coord): check rows_affected on the transition_session_created guard (finding #2 cont'd)

Completes the previous commit: the WHERE clause guard alone silently
swallowed a lost race (0 rows matched) as Ok(()) unless rows_affected()
is actually checked. This was split out of the prior commit by mistake
during hunk staging; closing the gap here.

* docs(coord): fix stale FOR UPDATE lock-duration comment (finding #3)

"Held only for the pick + insert below (sub-ms)" stopped being true once
reserve_and_persist_create's satellite writes (sealed-secrets insert,
per-capability insert loop, integration-policy upsert) moved inside the
same transaction as the FOR UPDATE host-row lock (issue #535 (b)) — the
lock is now held until tx.commit() at the end of the function, across
all of that. Correct the comment so the next reader doesn't under-estimate
placement-lock contention on a many-capability create burst.

* refactor(coord): delete dead upsert_session_secrets/set_session_harness (finding #4)

Both are zero-caller writers left behind by reserve_and_persist_create
subsuming the old satellite-write paths: the sealed-secrets bytea and the
harness selection now ride the one-transaction write-set / row INSERT
directly (engram-postgres/src/lib.rs's reserve_and_persist_create), not
these standalone upserts. Confirmed zero callers workspace-wide (including
orchestrator/ and web/) before deleting — only the trait declarations, the
PostgresStore impls, and 10 mock impls referenced them.

Per the repo's clean-break convention, retire both from the trait, the PG
impl, and every mock rather than leaving them as an orphaned, unused
re-entry point for the FK-ordering/partial-write bug class this PR set out
to kill. get_session_secrets/delete_session_secrets and get_session_harness
are untouched — those remain live (read/delete) call sites.

* fix(coord): split INSERT/UPDATE migration triggers to fix invalid WHEN-OLD DDL (finding #1)

migration 0077's `hosts_notify_fleet_catalog_changed` trigger's WHEN
clause referenced OLD on an AFTER INSERT OR UPDATE trigger. Postgres
rejects this at CREATE TRIGGER time ("INSERT trigger's WHEN condition
cannot reference OLD values") — OLD doesn't exist on INSERT and the
restriction is static, not runtime, so the `OLD IS NULL` guard didn't
help. This is the exact error CI hit ("while executing migration 77")
and would crash-loop every coordinator replica at boot on merge, since
migrations run at coordinator startup.

Split into two triggers: an INSERT trigger with no WHEN clause (a new
host's first bundle stamp always counts as a "change"), and an UPDATE
trigger with `WHEN (OLD.current_bundles IS DISTINCT FROM NEW.current_bundles)`.
Also updates the pg_listener.rs comment describing the guard now that
it only applies to the UPDATE leg.

* chore(migrations): renumber 0077 -> 0082 (batch land-queue collision)

Six PRs in this land batch each added a migration numbered 0077. Land-queue
assignment: #560 keeps 0077, #561->0078, #563->0079, #564->0080, #565->0081,
this PR (#566)->0082. Pure rename plus updating the two in-repo comments
that named the migration by number; no SQL content change.

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 7, 2026
… the lifecycle kernel (#543) (#599)

* feat(ops): ADR 0079 — durable per-session op log with fencing epochs: the lifecycle kernel (#543)

Every session lifecycle verb (resume, evict, deliver, create_boot,
destroy) is now a durable session_ops PG row driven by a single-writer-
per-session executor. sessions.current_epoch is the fencing epoch:
CAS-bumped in the same transaction that claims an op, appended
(AND current_epoch = $e) to every session-row write an op makes, and
carried (SessionFence) on every session-scoped host RPC, gated by the
host's persisted per-session high-water (work_dir/epochs, rejecting
stale with FAILED_PRECONDITION). WIRE_VERSION 11 -> 12 (lockstep roll).

The executor (session_ops.rs) is LISTEN/NOTIFY-hot (pg_notify
'session_ops'; 5s poll = fallback only), claims inline in the enqueue
transaction on the idle-session happy path (one PG round trip), records
durable per-step markers (idempotent-from-step crash resume), and
replaces the lease reaper with fence-then-resume reclaim. Manual
snapshot / evac resume / live teleport ride inline OpClaims on the same
primitive pending their own verb phases.

DELETED (grep-zero): the SessionLeaseGuard ecosystem + heartbeat +
reaper + LeaseTouch taxonomy, the session_lease table (migration 0093)
and its MetadataStore surface, both 8x3s lease-acquire retry loops, the
Evicting hold + its polls/env, the resume/evict/snapshot spawn-detach
pipelines, the residual-sandbox destroy compensation, the queue
scanner's requeue-by-poll, and the prompt path's mid-move HOLD. The
evict-then-resume collision class (3-21s stalls) is structurally gone:
a resume behind an in-flight evict is ordering by log, zero sleeps.

The deliver verb absorbs the outbox driver's per-session single-flight
(rows/ack/202 framing stay ADR 0073's) and performs the ADR 0074
rung-1/2 ascent inline under its own fence. Epoch-0 reject is DEFERRED
(allow-but-don't-advance) for the four named out-of-op senders — see
the ADR divergence log and check_session_epoch's disposition table.

Migrations 0092 (session_ops + current_epoch) + 0093 (drop
session_lease). Tests: session_ops_live_pg (claim CAS, one-running,
fencing 0-rows, ordering zero-sleeps, idempotency, cancellation,
backoff-yields-head), session_epochs host tests, verb-level ordering
units, migrated scanner/evicting-gate/outbox suites. Workspace 1501
passed; CI live-PG lane 106 passed; clippy/fmt/hakari clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx

* fix(ops): fence the evict/manual-snapshot record so a reclaimed-out op can't land a phantom recoverable row (ADR 0079 re-review #3/#4)

The `park_or_capture` eviction step brackets pause → snapshot → record →
commit with no intermediate `ctx.step()` fence check, and the capture leg
can run minutes. A coord↔PG partition that outlasts `RECLAIM_STALE` (180s)
lets a successor op re-claim the session (CAS-bumping `current_epoch`)
while the predecessor is mid-capture. `record_snapshot` was a PLAIN,
unfenced INSERT: on partition-heal the fenced-out predecessor would land a
`recoverable` row a resume could pick (the 89f7984d durability-lie class)
and then issue `commit_snapshot` under the stale epoch — the host's
per-session epoch high-water is only eventually-consistent with PG (a
reclaim bumps PG but not the host until the successor's first fenced RPC),
so that commit can slip through.

Add `MetadataStore::fenced_record_snapshot(snap, epoch)`: it writes the
row ONLY while `sessions.current_epoch == epoch`, atomically in one
transaction (the fence read is `FOR UPDATE`, serializing against the
claim/reclaim CAS), returning Ok(false) when fenced. The Postgres impl
routes both `record_snapshot` and the fenced variant through one private
`record_snapshot_guarded(snap, fence: Option<i64>)` so the INSERT +
generation bump + durable-head advance stay byte-identical; the default
trait impl delegates (fence-less) for the in-memory mocks. The eviction
pipeline and the manual-snapshot `recoverable=true` promote now use it and
bail cleanly (no phantom row, no `commit_snapshot`) when fenced. PG is the
authority, so this closes the window WITHOUT the proactive
host-epoch-advance-at-reclaim (ADR 0079 deferral #1c stays a pure
optimization). Live-PG regression:
`fenced_record_snapshot_writes_only_under_current_epoch`.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K

* fix(migration): heartbeat the teleport claim's synchronous body + fix stale RECLAIM_STALE comments (ADR 0079 re-review #1/#5)

`migrate_session_live` held its `OpClaim` across the entire synchronous
move body (presetup → concurrent restore-await → blackout → rebind →
reactivate) with NO liveness beat — the only `heartbeat_at` stamp was
`try_acquire`'s, and the first refresh (`claim.touch`) is in the finalize
drain loop, which runs AFTER the body. A move whose body outlives
`RECLAIM_STALE` (180s: a large VM over a slow inter-host link) was
therefore reclaimed out from under a PERFECTLY HEALTHY holder — no
partition required — and that holder then kept driving its (un-fenced,
#7) migration RPCs while the successor's reclaim failed the row: a
double-drive. The manual-snapshot inline claim already beats its body via
`spawn_heartbeat`; the teleport body now does too, dropped right before
the finalize task takes over its own beat.

Also fix two stale "60s staleness" comments (`RECLAIM_STALE` was raised
60s → 180s in review finding #1): the drain-loop touch comment in
live_migration.rs and the `OpClaim::spawn_heartbeat` doc in session_ops.rs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K

* fix(migration): fence the source-mutating migration RPCs so a stale teleport holder can't blackout/commit a live source (ADR 0079 re-review #7)

`migration_capture`/`_presetup`/`_capture_postcopy`/`_commit`/`_abort`
carried NO fencing_epoch, so `check_session_epoch` never gated them — the
one class of session-scoped host RPC the host could not reject. Combined
with the teleport claim's (now-fixed) missing heartbeat, a coord↔PG
partition (coord↔host is a SEPARATE gRPC transport, ADR 0013) that
outlasts RECLAIM_STALE lets a successor pod reclaim the teleport op while
the original holder stays alive and drives migration_capture_postcopy
(blackout) or migration_commit (destroys the source VM) with nothing able
to reject it.

Thread a SessionFence through all five: MigrationCapture/MigrationPresetup
now take FencedSandboxRequest; MigrationExportRef gains fencing_epoch +
session_id (MigrationCapturePostCopy/Commit/Abort). The host-agent gRPC
server runs check_session_epoch on each and forwards the fence to the
backend seam; the coordinator's teleport pipeline stamps claim.fence() at
every call site. MigrationFetch (host-to-host, export-nonce gated) and
MigrationDrainWait (dest-side) are deliberately unfenced — no
source-mutating write on the holder's behalf. SandboxBackend stays
fence-free (backend-local); only the HostClient wire seam carries it. No
WIRE_VERSION bump — the fields ride the unreleased v12 (clean break).

Also threads the fence through two Linux-gated FC teleport tests whose
`restore`/`migration_*` client calls the base ADR 0079 commit left
un-updated (macOS clippy can't see cfg(target_os="linux") — a
green-local/red-CI hole); they pass SessionFence::unfenced() (epoch 0 is
the host's allow-but-don't-advance interim).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K

* docs(adr-0079): record the post-merge adversarial re-review fixes

Adds a divergence-log section documenting the re-review of the two
deferrals: the shared "a reclaim only fires on a genuinely dead executor"
justification is false under a coord↔PG partition (ADR 0013 makes
coord→host a separate transport). re-#1 (teleport body heartbeat), re-#7
(migration-RPC fencing, was deferred), and re-#3/#4 (fenced
record_snapshot) are FIXED; re-#1c (proactive host-epoch-advance) stays
deferred but re-grounded as a pure optimization now that PG-side fenced
writes make the host high-water's eventual-consistency window
harmless-by-construction; re-#5 (stale comments) fixed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 18, 2026
…hase 2 bookend close

- detect-rebake-lanes.py: a `test_host_sim` flag on the release closure
  of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store
  ride in as normal deps; disjoint from engram-dst's coordinator closure
  by construction).
- ci.yml: the `test-host-sim` lane — a whole-binary replay-twice
  determinism self-check (one seed run twice, stdout diffed), then fixed
  windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host
  <n>` replays any failure), own rust-cache key
  (workspace-release-host-sim), added to `CI Gate.needs:` (never
  individually required — the aggregator rule).
- nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping
  windows (360 chaos + 120 calm x 4000 steps; sized below the
  coordinator swarm for the host sim's real-fs step cost),
  --failure-report -> the same slug-deduped sim-failure issue filing,
  title-prefixed "nightly host sim:" so host slugs never collide with
  same-named coordinator invariants.
- ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the
  full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory,
  and the #4-subsumed-by-P4.5 note.

Workflow YAML validated (yaml.safe_load + actionlint — remaining
findings are the pre-existing classes: custom Blacksmith runner labels
and the same shellcheck style infos the sibling jobs carry).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 18, 2026
…hase 2 bookend close

- detect-rebake-lanes.py: a `test_host_sim` flag on the release closure
  of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store
  ride in as normal deps; disjoint from engram-dst's coordinator closure
  by construction).
- ci.yml: the `test-host-sim` lane — a whole-binary replay-twice
  determinism self-check (one seed run twice, stdout diffed), then fixed
  windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host
  <n>` replays any failure), own rust-cache key
  (workspace-release-host-sim), added to `CI Gate.needs:` (never
  individually required — the aggregator rule).
- nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping
  windows (360 chaos + 120 calm x 4000 steps; sized below the
  coordinator swarm for the host sim's real-fs step cost),
  --failure-report -> the same slug-deduped sim-failure issue filing,
  title-prefixed "nightly host sim:" so host slugs never collide with
  same-named coordinator invariants.
- ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the
  full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory,
  and the #4-subsumed-by-P4.5 note.

Workflow YAML validated (yaml.safe_load + actionlint — remaining
findings are the pre-existing classes: custom Blacksmith runner labels
and the same shellcheck style infos the sibling jobs carry).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 18, 2026
…hase 2 bookend close

- detect-rebake-lanes.py: a `test_host_sim` flag on the release closure
  of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store
  ride in as normal deps; disjoint from engram-dst's coordinator closure
  by construction).
- ci.yml: the `test-host-sim` lane — a whole-binary replay-twice
  determinism self-check (one seed run twice, stdout diffed), then fixed
  windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host
  <n>` replays any failure), own rust-cache key
  (workspace-release-host-sim), added to `CI Gate.needs:` (never
  individually required — the aggregator rule).
- nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping
  windows (360 chaos + 120 calm x 4000 steps; sized below the
  coordinator swarm for the host sim's real-fs step cost),
  --failure-report -> the same slug-deduped sim-failure issue filing,
  title-prefixed "nightly host sim:" so host slugs never collide with
  same-named coordinator invariants.
- ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the
  full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory,
  and the #4-subsumed-by-P4.5 note.

Workflow YAML validated (yaml.safe_load + actionlint — remaining
findings are the pre-existing classes: custom Blacksmith runner labels
and the same shellcheck style infos the sibling jobs carry).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 18, 2026
…e model oracle (#786)

* ADR 0098 R2: the host-effect queue (in-flight interruption + message faults)

Pre-R2 every SimHostClient verb mutated world truth inline after one
optional delay, so RPC loss/reorder/duplication and a replica crash
BETWEEN the store commit (the ack the coordinator already holds) and the
host-side world effect were structurally impossible (the audit's finding
#2/#4).

A mutating verb now records its world mutation as a typed `Effect`.
Inline delivery is the default — byte-for-byte the pre-R2 world, so Calm
seeds are unchanged (Calm never opens a deferred window). When the target
host is in the deferred set (a Chaos fault window), the effect is queued
under a monotonic serial (BTreeMap → deterministic order) for a later
scheduler step to deliver / drop / duplicate / reorder:

- DeferHost(i, on)   — open/close a host's deferred window
- DeliverEffects     — deliver the queue in serial (causal) order
- DropEffect         — loss: drop one queued effect (seeded pick)
- DuplicateEffect    — re-deliver one effect (seeded pick)
- ReorderEffects     — deliver in a seeded-shuffled order

Crash/restart of a host severs its in-flight effects (never resurrected
onto the cleared VM set); quiescence closes every window and flushes the
queue in order before the fleet heals, so world truth is consistent with
the coordinator's commits. The Chaos weight table carves 9 points out of
AdvanceTime/Driver/HostHeartbeats/Crash/RestartHost for the new arms,
which shifts every Chaos seed's exploration (seeds pin to a commit); the
pinned chaos seeds still converge and are kept.

tests/effect_queue.rs pins the two reachable states the queue unlocks: a
deferred create withheld-then-delivered, and a replica crash inside the
commit→effect window followed by effect loss, both converging at
quiescence (no_op_dropped + no-stragglers + no-orphans).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ADR 0098 R2: API-driven workload over the real Router/gRPC surface

Pre-R2 the workload called store/core fns directly (0/17 HTTP, 0/~69 gRPC
driven — the audit's finding). This lands the real surface:

- workload.rs drives each replica's ACTUAL surface: the tonic
  AppSessionService/AppFleetService impls (auth + convert.rs + the same
  *_core the axum handler calls — handler-direct, no socket, to keep the
  paused-clock current-thread determinism) for create/prompt/resume/
  delete/drain, and the actual axum api::router via tower::oneshot
  (real-wire, middleware + extractors) for admin pause/resume and the
  harness-idle host ingestion. A module honesty table documents which
  verb takes which path.
- tests/api_surface.rs exercises it deterministically: create is acked,
  persisted (never lost), and REPLAYS byte-identically; prompt+delete
  round-trip; the axum router runs the real HTTP handlers to a normal
  response.

Real bug found + fixed: create_session's prepare path minted the session
id via a raw `SessionId::new()` (Uuid::new_v4) instead of the injected
`services.entropy` — an ADR 0098 D1 determinism leak (the id diverged
every replay; in prod OsEntropy makes this behavior-identical). The
API-create path now replays.

SimMeta gains four PG-faithful methods the real handlers reach
(previously panic-stubs): insert/get/delete_broker_token,
get/set_teleport_target, rebind_session_guarded,
list_active_assignments_with_budgets_on_host — with broker-token and
teleport-target conformance scenarios (ADR 0098 D4).

Scope note: folding this workload into the chaos/calm SWARM additionally
needs host-fidelity fixes (schedulable wire_version, staged ready_images)
that unmask a genuine, separate driver-double-boot + eviction-durability
class the pre-R2 sim silently suppressed by leaving hosts wire-skewed and
snapshots non-recoverable. Those are real findings for a follow-up; here
the surface is driven directly and deterministically rather than shipping
a red swarm lane.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* ADR 0098 R2: the expected-state model oracle (the auditor)

TigerBeetle's auditor shape: an expected-state model fed ONLY by ACKED
workload outcomes, diffed against world/SimMeta truth in the standing
invariant pass (every step + at quiescence). The honesty boundary is
explicit — because the model is fed only by acks, a lost-response op (a
create whose boot never durably established, an op whose reply dropped)
is legitimately absent from the model and never asserted on.

What it asserts:
- acked-create/resume never silently lost: a session the workload saw
  reach Active (durable row committed) still has a row on every later
  step, unless a later ACKED destroy retired it. A row vanishing under a
  live session is exactly the #570 symptom class (coordinator-unbind vs
  host-teardown-reconcile racing a session out of existence) and the
  durability-lie class this program exists to catch.
- read-your-acked-writes: the image a session was created with is never
  repainted (both replicas read the same shared SimMeta, so a present row
  is readable on either).

Fed from the op-path CreateSession/ResumeSession outcomes (record_if_live
records a session once observed Active). Wired into run()'s per-step and
quiescence checks alongside the existing oracles.

Non-vacuity is proven, not assumed: tests/model_oracle.rs drives a
session to Active, drops its row directly (the corruption the #570 race
produces), and asserts the auditor FIRES with `model-acked-create-not-
lost`; restoring the row clears it. `SimMetadataStore::with_db_mut` is the
test-only corruption injector.

Swarm (chaos 0..40 x600) stays green — the auditor is a safety invariant
the current code upholds; the canary test is what proves it can bite.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
nikhilunni added a commit that referenced this pull request Jul 21, 2026
…inator arm + op-quiescence oracle

The 2026-07-21 8174b7aa livelock (fixed in #824) lived in the one boundary
seam the cosim did not model: `world.heartbeat()` wrote host liveness
straight to the meta store, bypassing `host_http::heartbeat` where the
quarantined-survivor advertise → enqueue arm runs — so the
enqueue → fast-skip → re-enqueue loop was structurally invisible to the
sim even though it modeled both neighbors (the host's quarantine
classification AND the real evict verb).

- host_http: extract the advertise arm verbatim as
  `quarantined_survivor_advertise_core` (the run_once pattern the cosim's
  other coordinator seams already ride); the handler calls it.
- cosim: `Cosim::advertise_quarantined` builds the survivor set from the
  host's real `quarantined_unknown` × its binding table and drives the
  real core — prod's 5s heartbeat, co-simulated. `session_op_count` is
  the op-quiescence oracle read: under a fixed world state, repeated
  ticks must stop growing session_ops (unbounded growth is this class's
  signature — prod ran it to ~43k rows; the ADR 0093 423-row pileup is
  the same shape).
- tests/quarantine_advertise_livelock: the incident replayed at the
  boundary — the gap-A roll recipe leaves a quarantined survivor, the
  session is parked at Created (the ADR 0077 harness-failed park), then
  advertise/drive ticks must converge (first reap destroys the VM,
  clearing the advertise source; session settles HostLost) with ops
  quiescent. Fail-without verified: disabling the #824 guard arm
  reproduces the exact prod signature (advertised [1,1,1,1,1,1], op
  count growing every tick).

Deliberately NOT tested at the boundary: the ADR 0079 finding-#4
dedup-while-queued property — prod's `session_ops::enqueue` detaches an
inline drive, so a queued-but-unclaimed op is unreachable in a directed
harness; the property stays pinned by host_http's
`quarantined_survivor_readverts_dedup_to_one_op` (seeded running op).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01XhTVjDb59Khm2U5g5e8yY9
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file javascript Pull requests that update javascript code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant