Repository navigation
chore(web)(deps): bump framer-motion from 11.18.2 to 12.38.0 in /web - #4
Merged
nikhilunni merged 1 commit intoMay 10, 2026
Merged
Conversation
Contributor
|
@dependabot recreate |
dependabot
Bot
force-pushed
the
dependabot/npm_and_yarn/web/framer-motion-12.38.0
branch
from
May 10, 2026 22:39
a82fe27 to
9c021f4
Compare
Bumps [framer-motion](https://github.com/motiondivision/motion) from 11.18.2 to 12.38.0. - [Changelog](https://github.com/motiondivision/motion/blob/main/CHANGELOG.md) - [Commits](motiondivision/motion@v11.18.2...v12.38.0) --- updated-dependencies: - dependency-name: framer-motion dependency-version: 12.38.0 dependency-type: direct:production update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
dependabot
Bot
force-pushed
the
dependabot/npm_and_yarn/web/framer-motion-12.38.0
branch
from
May 10, 2026 22:45
9c021f4 to
5eaff14
Compare
nikhilunni
approved these changes
May 10, 2026
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
Tier 4 #3(a). PooledBackend gains `with_chunk_cache(cache)`; when present, `materialize_chunked_rootfs` routes chunk reads through the cache via `materialize_to_file_cached`. Chunks shared across manifests (canonical-base images, fork lineage) now serve from local NVMe on subsequent materializes instead of round-tripping BlobStorage every time. Wired in both binaries: - Coordinator `--mode=all`: cache root at `<local_path>/chunk-cache/`, default 200 GiB budget. - Standalone host-agent: same default at `<work_dir>/chunk-cache/`. HostAgent grows `with_chunk_cache(cache)` that forwards into PooledBackend during `run()`. A new `materialize_chunked_rootfs_uses_chunk_cache_when_present` test proves the cache path is exercised: after warming, the test deletes the underlying chunks from the blob store and asserts the second materialize still succeeds — it can only succeed if the chunks were served from the cache. Materialized-file orphan reap (3(b)) and chunk-store GC scheduler (#4) still pending — both pair naturally and need the same admin-endpoint + cron pattern. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
Tier 4 #4 — explicit-trigger chunk-store GC. Pairs with the project's "implicit-trigger features need a paired admin endpoint firing the same primitive" pattern (see disk-pressure / flush). Cron scheduler lands later; the admin endpoint is the testable + drain-before-redeploy ops shape. Surface: - POST /api/admin/gc-chunks?retain_secs=N (default 86400 = 24h) triggers a single GC pass. Returns GcChunksResult JSON with chunks_deleted, bytes_freed, chunks_retained_age, elapsed_ms, live_manifest_count, retain_secs. - New MetadataStore::list_live_disk_manifest_ids() returns DISTINCT manifest_ids referenced by any snapshot row. Backed by the idx_snapshots_disk_manifest partial index from migration 0018 — fast scan even on busy deployments. - Services.chunk_store: Arc-cheap clone of the same store the --mode=all PooledBackend already builds. Single source of truth per process. What the GC live-set DOES NOT yet cover: - enabled_images' canonical manifests (no snapshot row points at them until someone takes a snapshot using that image). Workaround: operators pass a generous retain_secs window so newly-baked images survive the GC gap. Real fix: persist disk_manifest_ref on enabled_images and union the two sets. Tracked in docs/chunked-storage-rollout.md. Verification: - 586/586 workspace nextest pass. - Two unit tests in tests/api.rs lock the empty-store response shape + the retain_secs query-param round-trip. - New live-PG test admin_gc_chunks_live_pg seeds a snapshot row with a manifest, fires the endpoint, asserts live_manifest_count reflects the seeded ref. CI runs it via the Postgres-gated step. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
Phase 5 follow-up #32 slice 1: plumb the per-image bake-time canonical memory manifest through the type system end-to-end. The image-builder doesn't yet auto-populate it (slice 2 of #32 implements the bake-time FC boot+pause+chunk dance), but the wiring is in place — once an operator pre-computes a canonical ref on a sidecar Linux+KVM job, the field flows cleanly from `bundle.json` → `CachedImage.bundle.canonical_memory_manifest` → `SandboxSpec.canonical_memory_manifest` → `FcSnapshotManifest. canonical_memory_manifest` → UFFD handler's `--canonical-manifest`. Plumbing: - `engram_core::types::sandbox::SandboxSpec.canonical_memory_manifest` (new, `#[serde(default)]` — pre-existing specs round-trip). - `engram_host_agent::image_cache::ImageBundle.canonical_memory_manifest` (new, serde-default). - `engram_image_builder::BuildRequest.canonical_memory_manifest` — optional pre-computed ref; written to `bundle.json` verbatim. - `engram_sandbox_firecracker::FcSnapshotManifest. canonical_memory_manifest` — now sourced from `spec.canonical_memory_manifest` at snapshot time (was hardcoded `None`). UFFD restore sees the canonical ref from the FC sidecar JSON the way it always has. - `PooledBackend::create()` lifts `cached.bundle.canonical_memory_manifest` onto `spec.canonical_memory_manifest` after image-cache resolve. Mass-edit: 18 `SandboxSpec { ... }` literals across the workspace (coord, sandbox-process, sandbox-vz, sandbox-firecracker tests, host-agent, protocol) gain `canonical_memory_manifest: None`. Same for 5 `BuildRequest` literals and 2 `ImageBundle` test constructors. ADR 0007 e2e test suite (`crates/engram-host-agent/tests/ adr_0007_e2e.rs`): - **#1 — snapshot path patches sidecar JSON's memory_manifest + chunks memory.bin into the store**: fake FC backend writes the on-disk shape a real FC snapshot produces; PooledBackend.snapshot chunks the emitted memory.bin and patches the JSON. Asserts metadata.memory_manifest is set, the JSON sidecar carries the matching ref, and `materialize_to_file` round-trips byte-for-byte. - **#2 — cross-host restore materialises memory.bin from chunks**: host A creates + snapshots; the local memory.bin is then deleted to simulate cross-host transfer (only state.bin + manifest.json staged); host B's PooledBackend (sharing the chunk store) calls restore — the wrap materialises memory.bin from the chunked memory_manifest before inner.restore sees it. Asserts the inner backend sees a fully-formed snapshot dir + bytes round-trip. - **#3 — trace_host_hint round-trips through sidecar JSON**: the snapshotting host's id flows through serde so a cross-host restoring backend can pass `--prefault-trace <hint>` to the UFFD handler. - **#4 — bundle.canonical_memory_manifest lifts onto SandboxSpec**: validates the wire shape so the canonical ref survives bundle → spec → snapshot manifest plumbing. - **#5 — content-addressed dedup across sessions**: same bytes chunked twice produce identical hashes (the property the cross-session dedup story relies on). - **#6 — Linux + nbd-module-gated NBD daemon spawn** (`#[ignore]`, validates the disk-side end-to-end on the dev VM): builds a disk manifest, attaches a real NBD daemon against `/dev/nbd0`, drops cleanly. Pure Rust, no FC, no KVM, no real kernel — runs in 0.42s on every CI lane (macOS + Linux). The Linux+KVM-gated tests for real FC + real NBD live in `crates/engram-sandbox-firecracker/ tests/snapshot_uffd.rs` + the planned `nbd_chunked_disk.rs` (task #34). 206 unit + integration tests pass workspace-wide: - engram-host-agent lib: 85 - engram-host-agent adr_0007_e2e: 5 - engram-sandbox-firecracker lib: 38 - engram-image-builder lib: 15 - engram-core lib: 25 - engram-coordinator lib: 43 - (the new test file adds 5 + 1 ignored) clippy + fmt clean on macOS; Linux check clean via dev VM.
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
- known-issues #9 (NBD) marked resolved with commit hashes - known-issues #10 (UFFD-from-chunks) marked resolved with the shipped surface enumerated (canonical capture, working-set R&R, cross-host materialize-from-chunks) - known-issues #13 (materialized-rootfs leak) marked resolved — reap_materialize_dir + admin endpoint + chunk_gc cron driver - rollout doc Phase 1 GC scheduler ⬜ → ✅ + Tier 4 #4 ditto - ADR 0007 "What this ADR does NOT cover" rewritten: NBD, UFFD, materialize-orphan-reap, 0018/0019 schema reshape all move from "not implemented" to shipped; observability + the Phase 6 destructive trait reshape remain the named gaps. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
Extends the Known Issues block with a "Pointers for the next session" subsection so a fresh session can pick up without re-deriving where to start. Per-issue: - entry-point file paths + grep anchors - proposed module additions (coord sweeper for #3, gauge for #4, heartbeat field for #5) - a SQL snippet for verifying stale templates against BlobStorage Plus: - suggested attack order (#3 first to unblock warm-path test, then #1+#2 as one commit, then #5, then #4 as defense) - ready-to-paste Cloud Logging queries for refill failures, snapshot dir leaks, eviction failures - production state snapshot at write time so a future reader can diff against current state (FC hosts pgs9+tf4j on 3b6aec3, demo image at warm-75babf7, harness still on pre-M1.12 b9dd2d1, 5+ known stale templates) - pointers to engrams-prod-ops skill scripts No code change — purely documentation hygiene for the handoff. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
ADR 0014 issue #4. Defense-in-depth backstop for the prod incident on `engrams-fc-xngk` (99 GB disk filled in ~13 min at ~25 leaked snapshot dirs × 4 GiB). Commits 1+2 (strict cascade + commit/abort lifecycle) already prevent the specific leak class that triggered the original incident, but if some future bug introduces a different on-disk leak, we want the host's idle-evict driver to stop adding fuel to the fire before disk fills. Mechanism: on every idle-evict tick (default 10s), the host-side driver in lib.rs calls `idle_evictor::disk_pressure_check(work_dir, floor)`. If statvfs reports free bytes below the floor, the tick: - increments `engram_host_idle_evict_disk_pressure_holds_total`, - logs a WARN with free / floor for ops to page on, - skips pushing candidates this tick. `disk_pressure_check` fails open on statvfs error — we'd rather over- evict than block all evictions silently on a transient FS hiccup. Default floor: 20 GiB. Sized for ~5 concurrent in-flight idle-evicts × ~4 GiB per FC memory dump + headroom (the actual `engrams-fc-xngk` post-mortem shows that 4 GiB-per-attempt at the unbounded rate fills the 99 GB disk in 13 min; 20 GiB free is the safety margin under which we stop pushing new candidates). Tunable via `ENGRAM_IDLE_EVICT_DISK_FLOOR_BYTES`. New host-agent metrics: - `engram_host_disk_free_bytes` (gauge) — sampled each tick. - `engram_host_idle_evict_disk_pressure_holds_total` (counter) — increments each tick we hold off. Tests added: - `free_disk_bytes_returns_value_for_tempdir` + `_none_for_missing_path` cover the statvfs probe's Ok/Err branches. - `disk_pressure_check_fails_open_on_statvfs_error` — missing path fails open (allow=true). - `disk_pressure_check_allows_when_floor_is_zero` and `disk_pressure_check_blocks_when_floor_exceeds_capacity` cover both sides of the gate. All 761 tests pass; `just check` clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 4, 2026
…(ADR 0037 P4d) Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off ⇒ every restore takes today's cold start_agent path; a warm snapshot's harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert in prod). - proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness (delivers session_id + session_env + first_prompt over the existing harness vsock channel). Threaded through the gRPC client/server, LocalHostClient (→ HarnessHub::bind), and HostRegistry routing. - coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND the base snapshot's warm_harness. When set, instead of start_agent: apply_egress_policy + late_bind_harness(session_env merged with the harness extras — forge/upload tokens, owner — + first prompt). The warm claude is not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_- PLACEHOLDER → real (reusing the token resolved for the cold per-session entry, swapping only the placeholder). warm_bind_enabled() kill-switch. KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell keep the capture-time (sentinel/placeholder) session env — FC merge_session_env is a documented no-op for the guest env. The agent loop is correct (forge/ upload tokens reach the helpers via the P4b session-env file; OAuth via the baked constant placeholder + proxy); a warm-path agentd env-merge RPC for shell/exec attribution is a follow-up. dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt clean. End-to-end FC warm-capture→restore→bind validation is next (P5). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 4, 2026
…ow-up Bookend update: P4b (no-respawn bind + token file), P4c (warm capture), P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e + measurement) and P6 (Accepted) remain — they need a real-FC warm path run (new hub-based e2e scaffolding) + a prod canary for the Accept-gate latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge gap (pitfall #4) as a follow-up. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 4, 2026
…(ADR 0037 P4d) Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off ⇒ every restore takes today's cold start_agent path; a warm snapshot's harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert in prod). - proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness (delivers session_id + session_env + first_prompt over the existing harness vsock channel). Threaded through the gRPC client/server, LocalHostClient (→ HarnessHub::bind), and HostRegistry routing. - coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND the base snapshot's warm_harness. When set, instead of start_agent: apply_egress_policy + late_bind_harness(session_env merged with the harness extras — forge/upload tokens, owner — + first prompt). The warm claude is not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_- PLACEHOLDER → real (reusing the token resolved for the cold per-session entry, swapping only the placeholder). warm_bind_enabled() kill-switch. KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell keep the capture-time (sentinel/placeholder) session env — FC merge_session_env is a documented no-op for the guest env. The agent loop is correct (forge/ upload tokens reach the helpers via the P4b session-env file; OAuth via the baked constant placeholder + proxy); a warm-path agentd env-merge RPC for shell/exec attribution is a follow-up. dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt clean. End-to-end FC warm-capture→restore→bind validation is next (P5). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 4, 2026
…ow-up Bookend update: P4b (no-respawn bind + token file), P4c (warm capture), P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e + measurement) and P6 (Accepted) remain — they need a real-FC warm path run (new hub-based e2e scaffolding) + a prod canary for the Accept-gate latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge gap (pitfall #4) as a follow-up. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 5, 2026
…(ADR 0037 P4d) Restore-side fork, gated coord-side by ENGRAM_WARM_HARNESS_BIND (default off ⇒ every restore takes today's cold start_agent path; a warm snapshot's harness is cleanly replaced by SpawnHarness's kill+respawn, so this is inert in prod). - proto/traits: new LateBindHarness RPC + HostClient::late_bind_harness (delivers session_id + session_env + first_prompt over the existing harness vsock channel). Threaded through the gRPC client/server, LocalHostClient (→ HarnessHub::bind), and HostRegistry routing. - coord create-flow fork (sessions.rs): warm_bind = warm_bind_enabled() AND the base snapshot's warm_harness. When set, instead of start_agent: apply_egress_policy + late_bind_harness(session_env merged with the harness extras — forge/upload tokens, owner — + first prompt). The warm claude is not respawned (V8 heap survives); its baked CONSTANT OAuth placeholder is what the proxy substitutes, so the egress policy maps WARM_CLAUDE_OAUTH_- PLACEHOLDER → real (reusing the token resolved for the cold per-session entry, swapping only the placeholder). warm_bind_enabled() kill-switch. KNOWN GAP (pitfall #4, flagged): on the warm path agentd's /exec + ttyd shell keep the capture-time (sentinel/placeholder) session env — FC merge_session_env is a documented no-op for the guest env. The agent loop is correct (forge/ upload tokens reach the helpers via the P4b session-env file; OAuth via the baked constant placeholder + proxy); a warm-path agentd env-merge RPC for shell/exec attribution is a follow-up. dev-vm: cargo check --workspace --all-targets + clippy -D warnings green; fmt clean. End-to-end FC warm-capture→restore→bind validation is next (P5). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 5, 2026
…ow-up Bookend update: P4b (no-respawn bind + token file), P4c (warm capture), P4d (restore-fork) all landed + gated off (inert in prod). P5 (FC e2e + measurement) and P6 (Accepted) remain — they need a real-FC warm path run (new hub-based e2e scaffolding) + a prod canary for the Accept-gate latency/density numbers. Flags the warm-path agentd /exec+ttyd env-merge gap (pitfall #4) as a follow-up. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 8, 2026
Define the K-phase labels (used across chart comments + PRs but never written down) and map them to the migration-path steps. Add design sections for K3 (HostFleet CRD + drain-gated operator), K4 (demand autoscaling off the coordinator capacity signal), and K5 (parallel-run cutover + node-asset staging + MIG retirement). Record the stable-HostId completion in K2; mark open questions 1-2 resolved by the K3/K4 designs and refine #4 (pod-netns settled; dedicated tainted pool + PSA namespace + KVM device plugin as the remaining security posture). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jun 25, 2026
…e leak) (#439) Each enabled-image re-bake/refresh captures a fresh per-image base snapshot and swaps `enabled_images.base_snapshot_id` to it (`upsert_enabled_image`'s ON CONFLICT), leaving the PRIOR base row dangling: `session_id IS NULL`, referenced by no `enabled_images` row. Nothing deleted it — `prune_session_snapshots` (checkpoint retention) is `session_id IS NOT NULL` only — and an orphan base keeps pinning its own disk+memory chunks via pin-set sources #3/#4 (`list_recoverable_snapshot_{disk,memory}_manifests`, which filter on `recoverable = TRUE` with no session predicate). So every image refresh permanently leaked one base snapshot's chunks (20-32 GB for the heavy dogfood images). Prod had 23 such orphans (~150 GB logical) going back 9 days, none in any GC queue. Add `prune_orphan_base_snapshots`, the `session_id IS NULL` mirror of `prune_session_snapshots`: delete base rows older than `ENGRAM_BASE_SNAPSHOT_RETENTION_HOURS` (default 24) that are referenced by no `enabled_images.base_snapshot_id` (live OR soft-deleted — soft-deleted lineage is intentionally still chunk-pinned, ADR 0021 P1.8), bumping `chunk_generation` in the same TX (GC-barrier symmetry). The existing chunk-GC (ADR 0016 Phase C) and snapshot-blob-GC (ADR 0028 addendum) sweeps then reclaim the now-unpinned chunks and portable `snapshots/<id>/` blobs — the reaper deletes nothing in BlobStorage directly. The `base_snapshot_id` FK (REFERENCES snapshots(id), no ON DELETE) is a hard backstop against ever deleting an in-use base. Spawned beside `checkpoint_retention`. A sweeper (not an inline delete-old-base in RefreshImage) so it both drains the existing backlog and survives restarts, per the "scanner drives transitions" convention. Also fixes the now-stale `snapshot_blob_pin_set` doc that claimed base rows are never deleted. Test (live-PG, wired into the existing Postgres-gated CI lane): a superseded orphan past grace is reaped; the current (enabled-image) base and a fresh orphan within grace are kept; session snapshots are untouched. The sibling checkpoint-retention test's template is pinned recent so the new global reaper can't collect it under local parallel runs (CI serializes the lane). Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 3, 2026
…ss (finding #4) Both are zero-caller writers left behind by reserve_and_persist_create subsuming the old satellite-write paths: the sealed-secrets bytea and the harness selection now ride the one-transaction write-set / row INSERT directly (engram-postgres/src/lib.rs's reserve_and_persist_create), not these standalone upserts. Confirmed zero callers workspace-wide (including orchestrator/ and web/) before deleting — only the trait declarations, the PostgresStore impls, and 10 mock impls referenced them. Per the repo's clean-break convention, retire both from the trait, the PG impl, and every mock rather than leaving them as an orphaned, unused re-entry point for the FK-ordering/partial-write bug class this PR set out to kill. get_session_secrets/delete_session_secrets and get_session_harness are untouched — those remain live (read/delete) call sites.
nikhilunni
added a commit
that referenced
this pull request
Jul 3, 2026
…licated ~60-line copy) PR #556 review finding #4 [CONFIRMED]. `api::prompt`'s emit-ordering tests carried a ~60-line copy of the exact `AppState`/`Services`/`MiniMeta` wiring `api::snapshot`'s `evicting_gate_tests` already had, justified by a comment claiming test modules "can't cheaply share private test-only fns" — false for same-crate unit tests, as `state::tests::MiniMeta` itself already demonstrates being shared across files. Hoists one `pub(crate) fn build_state_for_session` next to `MiniMeta` in `state::tests`; both call sites now delegate to it (`api::snapshot`'s copy trims to a 2-tuple wrapper since its tests don't need the `MiniMeta` handle). Every future `Services` field addition now has one call site to update, not two. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc
4 tasks done
nikhilunni
added a commit
that referenced
this pull request
Jul 6, 2026
… overlapped boot legs + prompt-over-wire (#566) * feat(coord): migration 0077 — session.selected_skills + fleet_catalog_changed NOTIFY trigger Part of issue #535 (create-as-a-plan). Groundwork for two later commits: - sessions.selected_skills (TEXT[]) persists a create's dynamic-mount selection so the queue scanner's boot re-prepare can reconstruct it (ADR 0055 TODO(P1-D): queued creates currently boot with base skills only, since the queue row never carried the selection). - notify_fleet_catalog_changed() + a trigger on hosts scoped to current_bundles changes (guarded by IS DISTINCT FROM, so it stays quiet across the few-seconds heartbeat UPDATE and only fires on an actual host-roll stamp change) backs the coordinator's boot-bundle cache invalidation, added in a follow-up commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): per-enabled-image boot bundle cache (issue #535 a) Combines the issue's steps 2+3 into one commit (the NOTIFY listener's invalidation calls need the cache type to exist first; splitting them would leave a dead intermediate compile state). - New `boot_bundle` module: `BootBundleCache` read-through caches, per enabled image, the parsed manifest + fetched base-snapshot record + resolved memory/vcpu budgets (today re-derived on every create), plus the fleet's baked bundle-name→sha catalog (today re-scanned via `list_active_hosts` up to three times per create). Both are TTL'd (30s belt-and-braces against a dropped PgListener notification). - `engram-postgres`: `upsert_enabled_image` / `soft_delete_enabled_image` now fire `pg_notify('enabled_image_changed', image_uri)` inside their existing transaction (delivered iff it commits); `delete_enabled_image` fires it best-effort after, mirroring `org_secret_changed`. - `pg_listener`: subscribes `enabled_image_changed` (invalidates one cache entry) and `fleet_catalog_changed` (invalidates the whole catalog — see migration 0077's trigger, landed in the prior commit). - `prepare_from_grpc` / `prepare_from_row` / `prepare_inner` / `fleet_bundle_catalog` rewired onto the cache: the strict (non-soft-deleted) vs. tolerant (`_any`) split moves to the two call sites (the cache always fills via the tolerant view), and `boot_on_reserved_host`'s per-create `get_snapshot` is gone — the snapshot record now rides `BootInputs.base_snapshot` from the bundle. Net: a warm-cache create now does zero `toml::from_str` calls and zero extra `list_active_hosts` scans beyond the one placement still needs (`candidates_for` — deliberately kept, per the issue's "conscious divergence": placement needs a heartbeat-fresh host view). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): one-transaction session write-set (issue #535 b) Collapses `reserve_placement` + `enqueue_session_create` + the boot/ enqueue paths' separate satellite-write chains into a single `MetadataStore::reserve_and_persist_create` call whose Postgres impl commits the ENTIRE write-set — the row (placed or queued) plus every satellite (sealed secrets, capabilities, integration policy, harness, selected skills) — in ONE `FOR UPDATE` transaction, before any host RPC. - `engram-core`: new `SessionCreateWriteSet` / `CreateDisposition` + `MetadataStore::reserve_and_persist_create` (replaces `reserve_ placement` + `enqueue_session_create`) and `transition_session_created` (a slim `pending → created` + `sandbox_id` UPDATE, replacing `create_ session_created`'s INSERT-or-UPDATE upsert — the row is now guaranteed to already exist). `Session` gains `selected_skills: Vec<String>`, fixing the ADR 0055 TODO(P1-D) gap: a queued create's boot re-prepare can now reconstruct its dynamic-mount selection instead of silently dropping to base skills. - `engram-postgres`: the transactional impl (extends `reserve_placement`'s FOR-UPDATE body); `get_session`/`list_active_sessions`/ `list_queued_sessions_fifo` project the new column. - `engram-coordinator`: `boot_prepared` seals secrets (KEK, pure crypto — has no place inside the DB transaction) and serializes the policy BEFORE calling `reserve_and_persist_create`, then dispatches on `CreateDisposition` — `enqueue_create` as a separate function is gone, its Queued-disposition handling folds into `boot_prepared`. `boot_on_reserved_host` now does exactly ONE write of its own (`transition_session_created`, since the sandbox doesn't exist until the restore RPC returns) — the FK-ordering bug class (a satellite write racing the row's own insert; the ADR 0051 forge-token regression) is dead by construction, not by "call it after the row" convention. - Every other `MetadataStore` impl (9 test/mock fixtures across 6 crates) updated: the 2 that exercise the real create path (coordinator HTTP + gRPC integration tests) got honest in-memory equivalents; the rest mirror their pre-existing `unreachable!()`/`unimplemented!()` convention for unexercised trait surface. - New Postgres-level tests (`placement_reservation_live_pg.rs`) proving the write-set's atomicity: a single call commits every satellite together, and a forced mid-transaction failure (duplicate session_id) leaves NOTHING from that attempt — not even satellites that would have followed the failing statement. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): overlap the independent boot legs with the restore RPC (issue #535 c) `boot_on_reserved_host` no longer serializes the restore RPC behind the env/egress work (or vice versa) — the two are independent (neither touches the other's inputs) and now run concurrently via `tokio::join!`: - Restore leg: `restore_base_on_host` (the ~0.4-0.7s VM-side work). - Env/egress leg: the per-spawn forge/upload broker-token mint (`inject_harness_env`) + the integration policy's Plane-B injection resolution (`resolve_inject_entries`, which can round-trip an external mint-provider API for a mint-mode connector) — this is where that external round trip moves OFF the serial tail. Both only need `session_id`/`image_ref`/`integration_policy`, not the sandbox; the broker-token FK has been satisfiable since `reserve_and_persist_ create` committed the row, well before this function runs. `build_egress_policy` splits accordingly: `resolve_inject_entries` + `build_observe_entries` (sandbox-independent, now called from the overlapped leg) stay as-is; the renamed `assemble_egress_policy` is the remaining sandbox-dependent half (`guest_ip` + final assembly). Also parallelizes `resolve_policy_secrets`' per-secret `SecretStore` round trips (order-insensitive — no secret depends on another) via `futures::future::join_all`, replacing the one-at-a-time loop. `transition_session_created` + the removal of `boot_on_reserved_host`'s satellite writes already landed in the prior commit (they're the same underlying change as the one-transaction write-set — splitting them would have left a dead intermediate compile state), so this commit is scoped to the actual leg-overlap + secret-resolution parallelization. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord,harness): prompt over the wire (issue #535 d) The initial prompt no longer rides `ENGRAM_INITIAL_PROMPT` env — every prompt, first or follow-up, is now a harness-protocol `Prompt` frame: - `api/prompt.rs`: factored the echo-then-forward core out of `send_prompt_core` into `deliver_prompt(state, session_id, sandbox_id, prompt_id, text)` — the user-echo-first ordering (load-bearing for web rendering) and the self-healing `deliver_with_reattach` forward, shared by every caller. - `session_boot::boot_on_reserved_host`: after the Active flip, mints a server-side `prompt_id` and calls `deliver_prompt` for the initial prompt — replacing the synthetic `prompt_id: None` event. Delivery failure past the reattach budget is `BootError::Started` (terminal, consistent with a `start_agent` failure): a session that can't receive the prompt that created it is broken. - `resolve_harness` no longer takes an `initial_prompt` param or inserts `ENGRAM_INITIAL_PROMPT`; `git grep ENGRAM_INITIAL_PROMPT` now returns nothing. - `engram-harness-claude`: `run_engine` drops the `initial_prompt` param — the pending queue starts empty and the first prompt arrives via `HarnessCommand::Prompt` like every other one. Updated the 8 unit tests that seeded an initial prompt through the deleted parameter to instead send it via `cmd_tx` post-spawn (and to expect the leading `Idle` the engine now emits before any prompt arrives, since the env-seeded fast-path — "start turn 1 with no leading Idle" — no longer exists). - `engram-host-agent/tests/e2e_harness.rs`: `capture_sink` now also returns a command sender so `drive_harness` can push the initial prompt as a wire frame instead of an env var — a hand-rolled minimal stand-in for `HarnessHub` (this test drives `SandboxBackend` directly, no coordinator/hub in the loop). The queued path unifies for free: `SessionCreateWriteSet::queue_prompt` (landed in the write-set commit) is already the durable prompt, and `prepare_from_row` threads it into the identical `boot_on_reserved_host` delivery path — no separate queued-prompt spelling. Deviation: did not extend `e2e_stack.rs`'s create-with-prompt test (`e2e_claude_with_bogus_key_surfaces_anthropic_auth_error`) to assert `prompt_id` threads onto `run_started` — it's quarantined (#403, excluded from the gating e2e lane) and requires a live `ENGRAM_E2E_GRPC_ADDR` stack this environment doesn't have, so the change is unverifiable here. The same property (RunStarted.prompt_id matches the delivered prompt_id) is covered by engram-harness-claude's `queue_holds_edits_and_consumes_type_ahead` unit test instead. Note: engram-harness-claude's test module is `#[cfg(target_os = "linux")]`-gated and e2e_harness.rs is Linux+KVM+FC+Docker+sudo-gated — neither compiles or runs on this macOS dev machine; both are verified by careful reading + will run for real in CI's Linux/FC lanes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): coord_prepare/coord_finalize phase metrics + doc pass (issue #535) Adds the two phase labels the issue's acceptance criteria need to turn "the coordinator serial tail is ~1s" from an estimate into a measurement, on the existing `engram_session_boot_seconds` histogram (no new metric names): - `coord_prepare`: `create_session_core` entry through `reserve_and_ persist_create`'s commit — the serial coordinator-side prefix ahead of the (now-concurrent, host-side) restore work. Recorded on the Placed path only. - `coord_finalize`: the restore RPC returning through the `created → active` transition — the coordinator-owned tail after the host hands back a live sandbox. Success path only. `total` minus (`coord_prepare` + `coord_finalize`) is the actual host-side restore RPC wall time — the split this issue's evidence section was missing. Doc pass: `session_boot.rs`'s module header now describes the (a)-(d) pipeline shape instead of the pre-refactor procedure; the FK-ordering guard comment in `prepare_inner` (the anchor the issue tracked as `sessions.rs:1197-1203`, drifted slightly by the time this landed) rewritten to describe the current dead-by-construction invariant instead of the historical hazard. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * fix(coord): guard transition_session_created on status=pending (finding #2) The UPDATE that flips a session row to `created` + binds `sandbox_id` had no status guard and ignored `rows_affected`, so it returned `Ok(())` even when the row was deleted (DeleteSession) or requeued (stale-pending scanner) while the restore RPC that precedes this call was in flight. That silently binds a live sandbox onto a gone/inconsistent row instead of hitting the existing `Err` teardown arm in session_boot.rs, which destroys the now-orphaned sandbox. Add `AND status = 'pending'` to the WHERE clause and return MetaError::NotFound on rows_affected() == 0, matching the convention used elsewhere in this file (e.g. delete_registry_credential). * fix(coord): check rows_affected on the transition_session_created guard (finding #2 cont'd) Completes the previous commit: the WHERE clause guard alone silently swallowed a lost race (0 rows matched) as Ok(()) unless rows_affected() is actually checked. This was split out of the prior commit by mistake during hunk staging; closing the gap here. * docs(coord): fix stale FOR UPDATE lock-duration comment (finding #3) "Held only for the pick + insert below (sub-ms)" stopped being true once reserve_and_persist_create's satellite writes (sealed-secrets insert, per-capability insert loop, integration-policy upsert) moved inside the same transaction as the FOR UPDATE host-row lock (issue #535 (b)) — the lock is now held until tx.commit() at the end of the function, across all of that. Correct the comment so the next reader doesn't under-estimate placement-lock contention on a many-capability create burst. * refactor(coord): delete dead upsert_session_secrets/set_session_harness (finding #4) Both are zero-caller writers left behind by reserve_and_persist_create subsuming the old satellite-write paths: the sealed-secrets bytea and the harness selection now ride the one-transaction write-set / row INSERT directly (engram-postgres/src/lib.rs's reserve_and_persist_create), not these standalone upserts. Confirmed zero callers workspace-wide (including orchestrator/ and web/) before deleting — only the trait declarations, the PostgresStore impls, and 10 mock impls referenced them. Per the repo's clean-break convention, retire both from the trait, the PG impl, and every mock rather than leaving them as an orphaned, unused re-entry point for the FK-ordering/partial-write bug class this PR set out to kill. get_session_secrets/delete_session_secrets and get_session_harness are untouched — those remain live (read/delete) call sites. * fix(coord): split INSERT/UPDATE migration triggers to fix invalid WHEN-OLD DDL (finding #1) migration 0077's `hosts_notify_fleet_catalog_changed` trigger's WHEN clause referenced OLD on an AFTER INSERT OR UPDATE trigger. Postgres rejects this at CREATE TRIGGER time ("INSERT trigger's WHEN condition cannot reference OLD values") — OLD doesn't exist on INSERT and the restriction is static, not runtime, so the `OLD IS NULL` guard didn't help. This is the exact error CI hit ("while executing migration 77") and would crash-loop every coordinator replica at boot on merge, since migrations run at coordinator startup. Split into two triggers: an INSERT trigger with no WHEN clause (a new host's first bundle stamp always counts as a "change"), and an UPDATE trigger with `WHEN (OLD.current_bundles IS DISTINCT FROM NEW.current_bundles)`. Also updates the pg_listener.rs comment describing the guard now that it only applies to the UPDATE leg. * chore(migrations): renumber 0077 -> 0082 (batch land-queue collision) Six PRs in this land batch each added a migration numbered 0077. Land-queue assignment: #560 keeps 0077, #561->0078, #563->0079, #564->0080, #565->0081, this PR (#566)->0082. Pure rename plus updating the two in-repo comments that named the migration by number; no SQL content change. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 7, 2026
… the lifecycle kernel (#543) (#599) * feat(ops): ADR 0079 — durable per-session op log with fencing epochs: the lifecycle kernel (#543) Every session lifecycle verb (resume, evict, deliver, create_boot, destroy) is now a durable session_ops PG row driven by a single-writer- per-session executor. sessions.current_epoch is the fencing epoch: CAS-bumped in the same transaction that claims an op, appended (AND current_epoch = $e) to every session-row write an op makes, and carried (SessionFence) on every session-scoped host RPC, gated by the host's persisted per-session high-water (work_dir/epochs, rejecting stale with FAILED_PRECONDITION). WIRE_VERSION 11 -> 12 (lockstep roll). The executor (session_ops.rs) is LISTEN/NOTIFY-hot (pg_notify 'session_ops'; 5s poll = fallback only), claims inline in the enqueue transaction on the idle-session happy path (one PG round trip), records durable per-step markers (idempotent-from-step crash resume), and replaces the lease reaper with fence-then-resume reclaim. Manual snapshot / evac resume / live teleport ride inline OpClaims on the same primitive pending their own verb phases. DELETED (grep-zero): the SessionLeaseGuard ecosystem + heartbeat + reaper + LeaseTouch taxonomy, the session_lease table (migration 0093) and its MetadataStore surface, both 8x3s lease-acquire retry loops, the Evicting hold + its polls/env, the resume/evict/snapshot spawn-detach pipelines, the residual-sandbox destroy compensation, the queue scanner's requeue-by-poll, and the prompt path's mid-move HOLD. The evict-then-resume collision class (3-21s stalls) is structurally gone: a resume behind an in-flight evict is ordering by log, zero sleeps. The deliver verb absorbs the outbox driver's per-session single-flight (rows/ack/202 framing stay ADR 0073's) and performs the ADR 0074 rung-1/2 ascent inline under its own fence. Epoch-0 reject is DEFERRED (allow-but-don't-advance) for the four named out-of-op senders — see the ADR divergence log and check_session_epoch's disposition table. Migrations 0092 (session_ops + current_epoch) + 0093 (drop session_lease). Tests: session_ops_live_pg (claim CAS, one-running, fencing 0-rows, ordering zero-sleeps, idempotency, cancellation, backoff-yields-head), session_epochs host tests, verb-level ordering units, migrated scanner/evicting-gate/outbox suites. Workspace 1501 passed; CI live-PG lane 106 passed; clippy/fmt/hakari clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx * fix(ops): fence the evict/manual-snapshot record so a reclaimed-out op can't land a phantom recoverable row (ADR 0079 re-review #3/#4) The `park_or_capture` eviction step brackets pause → snapshot → record → commit with no intermediate `ctx.step()` fence check, and the capture leg can run minutes. A coord↔PG partition that outlasts `RECLAIM_STALE` (180s) lets a successor op re-claim the session (CAS-bumping `current_epoch`) while the predecessor is mid-capture. `record_snapshot` was a PLAIN, unfenced INSERT: on partition-heal the fenced-out predecessor would land a `recoverable` row a resume could pick (the 89f7984d durability-lie class) and then issue `commit_snapshot` under the stale epoch — the host's per-session epoch high-water is only eventually-consistent with PG (a reclaim bumps PG but not the host until the successor's first fenced RPC), so that commit can slip through. Add `MetadataStore::fenced_record_snapshot(snap, epoch)`: it writes the row ONLY while `sessions.current_epoch == epoch`, atomically in one transaction (the fence read is `FOR UPDATE`, serializing against the claim/reclaim CAS), returning Ok(false) when fenced. The Postgres impl routes both `record_snapshot` and the fenced variant through one private `record_snapshot_guarded(snap, fence: Option<i64>)` so the INSERT + generation bump + durable-head advance stay byte-identical; the default trait impl delegates (fence-less) for the in-memory mocks. The eviction pipeline and the manual-snapshot `recoverable=true` promote now use it and bail cleanly (no phantom row, no `commit_snapshot`) when fenced. PG is the authority, so this closes the window WITHOUT the proactive host-epoch-advance-at-reclaim (ADR 0079 deferral #1c stays a pure optimization). Live-PG regression: `fenced_record_snapshot_writes_only_under_current_epoch`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * fix(migration): heartbeat the teleport claim's synchronous body + fix stale RECLAIM_STALE comments (ADR 0079 re-review #1/#5) `migrate_session_live` held its `OpClaim` across the entire synchronous move body (presetup → concurrent restore-await → blackout → rebind → reactivate) with NO liveness beat — the only `heartbeat_at` stamp was `try_acquire`'s, and the first refresh (`claim.touch`) is in the finalize drain loop, which runs AFTER the body. A move whose body outlives `RECLAIM_STALE` (180s: a large VM over a slow inter-host link) was therefore reclaimed out from under a PERFECTLY HEALTHY holder — no partition required — and that holder then kept driving its (un-fenced, #7) migration RPCs while the successor's reclaim failed the row: a double-drive. The manual-snapshot inline claim already beats its body via `spawn_heartbeat`; the teleport body now does too, dropped right before the finalize task takes over its own beat. Also fix two stale "60s staleness" comments (`RECLAIM_STALE` was raised 60s → 180s in review finding #1): the drain-loop touch comment in live_migration.rs and the `OpClaim::spawn_heartbeat` doc in session_ops.rs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * fix(migration): fence the source-mutating migration RPCs so a stale teleport holder can't blackout/commit a live source (ADR 0079 re-review #7) `migration_capture`/`_presetup`/`_capture_postcopy`/`_commit`/`_abort` carried NO fencing_epoch, so `check_session_epoch` never gated them — the one class of session-scoped host RPC the host could not reject. Combined with the teleport claim's (now-fixed) missing heartbeat, a coord↔PG partition (coord↔host is a SEPARATE gRPC transport, ADR 0013) that outlasts RECLAIM_STALE lets a successor pod reclaim the teleport op while the original holder stays alive and drives migration_capture_postcopy (blackout) or migration_commit (destroys the source VM) with nothing able to reject it. Thread a SessionFence through all five: MigrationCapture/MigrationPresetup now take FencedSandboxRequest; MigrationExportRef gains fencing_epoch + session_id (MigrationCapturePostCopy/Commit/Abort). The host-agent gRPC server runs check_session_epoch on each and forwards the fence to the backend seam; the coordinator's teleport pipeline stamps claim.fence() at every call site. MigrationFetch (host-to-host, export-nonce gated) and MigrationDrainWait (dest-side) are deliberately unfenced — no source-mutating write on the holder's behalf. SandboxBackend stays fence-free (backend-local); only the HostClient wire seam carries it. No WIRE_VERSION bump — the fields ride the unreleased v12 (clean break). Also threads the fence through two Linux-gated FC teleport tests whose `restore`/`migration_*` client calls the base ADR 0079 commit left un-updated (macOS clippy can't see cfg(target_os="linux") — a green-local/red-CI hole); they pass SessionFence::unfenced() (epoch 0 is the host's allow-but-don't-advance interim). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * docs(adr-0079): record the post-merge adversarial re-review fixes Adds a divergence-log section documenting the re-review of the two deferrals: the shared "a reclaim only fires on a genuinely dead executor" justification is false under a coord↔PG partition (ADR 0013 makes coord→host a separate transport). re-#1 (teleport body heartbeat), re-#7 (migration-RPC fencing, was deferred), and re-#3/#4 (fenced record_snapshot) are FIXED; re-#1c (proactive host-epoch-advance) stays deferred but re-grounded as a pure optimization now that PG-side fenced writes make the host high-water's eventual-consistency window harmless-by-construction; re-#5 (stale comments) fixed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 17, 2026
Merged
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
…hase 2 bookend close - detect-rebake-lanes.py: a `test_host_sim` flag on the release closure of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store ride in as normal deps; disjoint from engram-dst's coordinator closure by construction). - ci.yml: the `test-host-sim` lane — a whole-binary replay-twice determinism self-check (one seed run twice, stdout diffed), then fixed windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host <n>` replays any failure), own rust-cache key (workspace-release-host-sim), added to `CI Gate.needs:` (never individually required — the aggregator rule). - nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping windows (360 chaos + 120 calm x 4000 steps; sized below the coordinator swarm for the host sim's real-fs step cost), --failure-report -> the same slug-deduped sim-failure issue filing, title-prefixed "nightly host sim:" so host slugs never collide with same-named coordinator invariants. - ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory, and the #4-subsumed-by-P4.5 note. Workflow YAML validated (yaml.safe_load + actionlint — remaining findings are the pre-existing classes: custom Blacksmith runner labels and the same shellcheck style infos the sibling jobs carry). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
…hase 2 bookend close - detect-rebake-lanes.py: a `test_host_sim` flag on the release closure of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store ride in as normal deps; disjoint from engram-dst's coordinator closure by construction). - ci.yml: the `test-host-sim` lane — a whole-binary replay-twice determinism self-check (one seed run twice, stdout diffed), then fixed windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host <n>` replays any failure), own rust-cache key (workspace-release-host-sim), added to `CI Gate.needs:` (never individually required — the aggregator rule). - nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping windows (360 chaos + 120 calm x 4000 steps; sized below the coordinator swarm for the host sim's real-fs step cost), --failure-report -> the same slug-deduped sim-failure issue filing, title-prefixed "nightly host sim:" so host slugs never collide with same-named coordinator invariants. - ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory, and the #4-subsumed-by-P4.5 note. Workflow YAML validated (yaml.safe_load + actionlint — remaining findings are the pre-existing classes: custom Blacksmith runner labels and the same shellcheck style infos the sibling jobs carry). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
…hase 2 bookend close - detect-rebake-lanes.py: a `test_host_sim` flag on the release closure of engram-dst-host itself (host-agent/host-core/engram-sim/chunk-store ride in as normal deps; disjoint from engram-dst's coordinator closure by construction). - ci.yml: the `test-host-sim` lane — a whole-binary replay-twice determinism self-check (one seed run twice, stdout diffed), then fixed windows (chaos 0..60 x 1000 steps, calm 0..30 x 1000; `just sim-host <n>` replays any failure), own rust-cache key (workspace-release-host-sim), added to `CI Gate.needs:` (never individually required — the aggregator rule). - nightly-sim.yml: the `host-swarm` job — date-derived non-overlapping windows (360 chaos + 120 calm x 4000 steps; sized below the coordinator swarm for the host sim's real-fs step cost), --failure-report -> the same slug-deduped sim-failure issue filing, title-prefixed "nightly host sim:" so host slugs never collide with same-named coordinator invariants. - ADR: P9 row -> Landed; the Phase 2 header flipped to COMPLETE with the full commit chain (P0-P9 + P4.5 + G1/G2), the flow/oracle inventory, and the #4-subsumed-by-P4.5 note. Workflow YAML validated (yaml.safe_load + actionlint — remaining findings are the pre-existing classes: custom Blacksmith runner labels and the same shellcheck style infos the sibling jobs carry). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 18, 2026
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
…e model oracle (#786) * ADR 0098 R2: the host-effect queue (in-flight interruption + message faults) Pre-R2 every SimHostClient verb mutated world truth inline after one optional delay, so RPC loss/reorder/duplication and a replica crash BETWEEN the store commit (the ack the coordinator already holds) and the host-side world effect were structurally impossible (the audit's finding #2/#4). A mutating verb now records its world mutation as a typed `Effect`. Inline delivery is the default — byte-for-byte the pre-R2 world, so Calm seeds are unchanged (Calm never opens a deferred window). When the target host is in the deferred set (a Chaos fault window), the effect is queued under a monotonic serial (BTreeMap → deterministic order) for a later scheduler step to deliver / drop / duplicate / reorder: - DeferHost(i, on) — open/close a host's deferred window - DeliverEffects — deliver the queue in serial (causal) order - DropEffect — loss: drop one queued effect (seeded pick) - DuplicateEffect — re-deliver one effect (seeded pick) - ReorderEffects — deliver in a seeded-shuffled order Crash/restart of a host severs its in-flight effects (never resurrected onto the cleared VM set); quiescence closes every window and flushes the queue in order before the fleet heals, so world truth is consistent with the coordinator's commits. The Chaos weight table carves 9 points out of AdvanceTime/Driver/HostHeartbeats/Crash/RestartHost for the new arms, which shifts every Chaos seed's exploration (seeds pin to a commit); the pinned chaos seeds still converge and are kept. tests/effect_queue.rs pins the two reachable states the queue unlocks: a deferred create withheld-then-delivered, and a replica crash inside the commit→effect window followed by effect loss, both converging at quiescence (no_op_dropped + no-stragglers + no-orphans). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ADR 0098 R2: API-driven workload over the real Router/gRPC surface Pre-R2 the workload called store/core fns directly (0/17 HTTP, 0/~69 gRPC driven — the audit's finding). This lands the real surface: - workload.rs drives each replica's ACTUAL surface: the tonic AppSessionService/AppFleetService impls (auth + convert.rs + the same *_core the axum handler calls — handler-direct, no socket, to keep the paused-clock current-thread determinism) for create/prompt/resume/ delete/drain, and the actual axum api::router via tower::oneshot (real-wire, middleware + extractors) for admin pause/resume and the harness-idle host ingestion. A module honesty table documents which verb takes which path. - tests/api_surface.rs exercises it deterministically: create is acked, persisted (never lost), and REPLAYS byte-identically; prompt+delete round-trip; the axum router runs the real HTTP handlers to a normal response. Real bug found + fixed: create_session's prepare path minted the session id via a raw `SessionId::new()` (Uuid::new_v4) instead of the injected `services.entropy` — an ADR 0098 D1 determinism leak (the id diverged every replay; in prod OsEntropy makes this behavior-identical). The API-create path now replays. SimMeta gains four PG-faithful methods the real handlers reach (previously panic-stubs): insert/get/delete_broker_token, get/set_teleport_target, rebind_session_guarded, list_active_assignments_with_budgets_on_host — with broker-token and teleport-target conformance scenarios (ADR 0098 D4). Scope note: folding this workload into the chaos/calm SWARM additionally needs host-fidelity fixes (schedulable wire_version, staged ready_images) that unmask a genuine, separate driver-double-boot + eviction-durability class the pre-R2 sim silently suppressed by leaving hosts wire-skewed and snapshots non-recoverable. Those are real findings for a follow-up; here the surface is driven directly and deterministically rather than shipping a red swarm lane. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ADR 0098 R2: the expected-state model oracle (the auditor) TigerBeetle's auditor shape: an expected-state model fed ONLY by ACKED workload outcomes, diffed against world/SimMeta truth in the standing invariant pass (every step + at quiescence). The honesty boundary is explicit — because the model is fed only by acks, a lost-response op (a create whose boot never durably established, an op whose reply dropped) is legitimately absent from the model and never asserted on. What it asserts: - acked-create/resume never silently lost: a session the workload saw reach Active (durable row committed) still has a row on every later step, unless a later ACKED destroy retired it. A row vanishing under a live session is exactly the #570 symptom class (coordinator-unbind vs host-teardown-reconcile racing a session out of existence) and the durability-lie class this program exists to catch. - read-your-acked-writes: the image a session was created with is never repainted (both replicas read the same shared SimMeta, so a present row is readable on either). Fed from the op-path CreateSession/ResumeSession outcomes (record_if_live records a session once observed Active). Wired into run()'s per-step and quiescence checks alongside the existing oracles. Non-vacuity is proven, not assumed: tests/model_oracle.rs drives a session to Active, drops its row directly (the corruption the #570 race produces), and asserts the auditor FIRES with `model-acked-create-not- lost`; restoring the row clears it. `SimMetadataStore::with_db_mut` is the test-only corruption injector. Swarm (chaos 0..40 x600) stays green — the auditor is a safety invariant the current code upholds; the canary test is what proves it can bite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 21, 2026
…inator arm + op-quiescence oracle The 2026-07-21 8174b7aa livelock (fixed in #824) lived in the one boundary seam the cosim did not model: `world.heartbeat()` wrote host liveness straight to the meta store, bypassing `host_http::heartbeat` where the quarantined-survivor advertise → enqueue arm runs — so the enqueue → fast-skip → re-enqueue loop was structurally invisible to the sim even though it modeled both neighbors (the host's quarantine classification AND the real evict verb). - host_http: extract the advertise arm verbatim as `quarantined_survivor_advertise_core` (the run_once pattern the cosim's other coordinator seams already ride); the handler calls it. - cosim: `Cosim::advertise_quarantined` builds the survivor set from the host's real `quarantined_unknown` × its binding table and drives the real core — prod's 5s heartbeat, co-simulated. `session_op_count` is the op-quiescence oracle read: under a fixed world state, repeated ticks must stop growing session_ops (unbounded growth is this class's signature — prod ran it to ~43k rows; the ADR 0093 423-row pileup is the same shape). - tests/quarantine_advertise_livelock: the incident replayed at the boundary — the gap-A roll recipe leaves a quarantined survivor, the session is parked at Created (the ADR 0077 harness-failed park), then advertise/drive ticks must converge (first reap destroys the VM, clearing the advertise source; session settles HostLost) with ops quiescent. Fail-without verified: disabling the #824 guard arm reproduces the exact prod signature (advertised [1,1,1,1,1,1], op count growing every tick). Deliberately NOT tested at the boundary: the ADR 0079 finding-#4 dedup-while-queued property — prod's `session_ops::enqueue` detaches an inline drive, so a queued-but-unclaimed op is unreachable in a directed harness; the property stays pinned by host_http's `quarantined_survivor_readverts_dedup_to_one_op` (seeded running op). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XhTVjDb59Khm2U5g5e8yY9
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps framer-motion from 11.18.2 to 12.38.0.
Changelog
Sourced from framer-motion's changelog.
... (truncated)
Commits
0bfc9fev12.38.0343cb0cUpdating layoutAnchoree99ad2Updating changelog062660bUpdating changgelog303da7dUpdating readmeb075adcMerge pull request #3647 from motiondivision/feat/layout-anchorf0991d6Add missing layoutAnchor !== false guard in attemptToResolveRelativeTargetb5798e9Merge pull request #3642 from motiondivision/worktree-fix-issue-30787686c19Merge pull request #3636 from motiondivision/worktree-fix-issue-3061a95c487Fix auto-scroll in reorder-virtualized test page