Repository navigation
chore(docker): bump rust from 1.83-slim to 1.95-slim in /docker in the docker-base group across 1 directory - #2
Merged
Conversation
Contributor
|
@dependabot recreate |
Bumps the docker-base group with 1 update in the /docker directory: rust. Updates `rust` from 1.83-slim to 1.95-slim --- updated-dependencies: - dependency-name: rust dependency-version: 1.95-slim dependency-type: direct:production dependency-group: docker-base ... Signed-off-by: dependabot[bot] <support@github.com>
dependabot
Bot
force-pushed
the
dependabot/docker/docker/docker-base-a5631364b5
branch
from
May 10, 2026 22:38
1fa205d to
9aa1445
Compare
nikhilunni
enabled auto-merge (squash)
May 10, 2026 22:51
nikhilunni
approved these changes
May 10, 2026
nikhilunni
added a commit
that referenced
this pull request
May 10, 2026
fetch-metadata's `update-type` output is null for grouped Dependabot PRs (even single-item groups like docker-base), which made the prior if-condition skip approval on PR #2 and similar. Walk `updated-dependencies-json` instead — auto-merge iff every dep in the array is patch or minor. Single-dep PRs end up with a 1-entry array, so the same code path covers both cases without branching. Also moves the update-type expression into env to dodge expression-in-shell interpolation risk in the approval body. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
… hosts Tier 4 #2 — real OCI credentials on the standalone host-agent. The infrastructure existed (engram-oci-auth::PgAuthResolver reads encrypted RegistryCredential rows from Postgres + decrypts under KEK, handles Static / GcpWorkloadIdentity / Anonymous) but only the coord could call it: --mode=all shared the OciClient with the in-process host-agent, but the standalone engram-host-agent binary fell back to AnonymousResolver — which meant zero private-registry support outside --mode=all. Now: the host-agent's OciClient is built around a `WsAuthResolver` that issues `RequestKind::ResolveRegistryAuth { host }` over the existing dialer WebSocket. Coord side installs an `AuthRequestHandler` that delegates straight to its PgAuthResolver. Plaintext creds traverse the WS only at pull time; never persisted on the host. Concrete changes: - engram-protocol: `RequestKind::ResolveRegistryAuth` + `ResponseKind::RegistryAuth { creds: Option<RegistryCreds> }`. New `HostRequestHandler` trait + handler hook on `ConnectedHost`. `HostSession` grows a request/response demuxer mirroring ConnectedHost so the host can issue RPCs over the same WS. - WIRE_VERSION bumped to 2: new enum variants shift bincode discriminants. The version handshake from 0df3a31 catches the mismatch loudly. - engram-coordinator: `Services.auth_resolver` surfaced as an Arc so `AuthRequestHandler` in api/hosts.rs can call it directly. Test fixtures updated to plant an `AnonymousResolver` there. - engram-host-agent: new `ws_auth` module with `WsAuthResolver` + `SessionHandle`. Dialer publishes the live `HostSession` into the handle on connect, clears on disconnect. `HostAgent` grows `with_auth_session_handle`; main.rs replaces the `AnonymousResolver` with the `WsAuthResolver` pair. - Rollout doc Tier 4 #2 reframed from "design call needed" to the concrete shape that landed. (The design call wasn't actually open — the auth pieces existed; the only gap was the WS hop.) Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
Cite ad13dc0 + summarize what landed: WsAuthResolver in the host-agent, AuthRequestHandler in the coord, WIRE_VERSION bump to v2 for the new RequestKind variant. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 12, 2026
Phase 5 follow-up #32 slice 1: plumb the per-image bake-time canonical memory manifest through the type system end-to-end. The image-builder doesn't yet auto-populate it (slice 2 of #32 implements the bake-time FC boot+pause+chunk dance), but the wiring is in place — once an operator pre-computes a canonical ref on a sidecar Linux+KVM job, the field flows cleanly from `bundle.json` → `CachedImage.bundle.canonical_memory_manifest` → `SandboxSpec.canonical_memory_manifest` → `FcSnapshotManifest. canonical_memory_manifest` → UFFD handler's `--canonical-manifest`. Plumbing: - `engram_core::types::sandbox::SandboxSpec.canonical_memory_manifest` (new, `#[serde(default)]` — pre-existing specs round-trip). - `engram_host_agent::image_cache::ImageBundle.canonical_memory_manifest` (new, serde-default). - `engram_image_builder::BuildRequest.canonical_memory_manifest` — optional pre-computed ref; written to `bundle.json` verbatim. - `engram_sandbox_firecracker::FcSnapshotManifest. canonical_memory_manifest` — now sourced from `spec.canonical_memory_manifest` at snapshot time (was hardcoded `None`). UFFD restore sees the canonical ref from the FC sidecar JSON the way it always has. - `PooledBackend::create()` lifts `cached.bundle.canonical_memory_manifest` onto `spec.canonical_memory_manifest` after image-cache resolve. Mass-edit: 18 `SandboxSpec { ... }` literals across the workspace (coord, sandbox-process, sandbox-vz, sandbox-firecracker tests, host-agent, protocol) gain `canonical_memory_manifest: None`. Same for 5 `BuildRequest` literals and 2 `ImageBundle` test constructors. ADR 0007 e2e test suite (`crates/engram-host-agent/tests/ adr_0007_e2e.rs`): - **#1 — snapshot path patches sidecar JSON's memory_manifest + chunks memory.bin into the store**: fake FC backend writes the on-disk shape a real FC snapshot produces; PooledBackend.snapshot chunks the emitted memory.bin and patches the JSON. Asserts metadata.memory_manifest is set, the JSON sidecar carries the matching ref, and `materialize_to_file` round-trips byte-for-byte. - **#2 — cross-host restore materialises memory.bin from chunks**: host A creates + snapshots; the local memory.bin is then deleted to simulate cross-host transfer (only state.bin + manifest.json staged); host B's PooledBackend (sharing the chunk store) calls restore — the wrap materialises memory.bin from the chunked memory_manifest before inner.restore sees it. Asserts the inner backend sees a fully-formed snapshot dir + bytes round-trip. - **#3 — trace_host_hint round-trips through sidecar JSON**: the snapshotting host's id flows through serde so a cross-host restoring backend can pass `--prefault-trace <hint>` to the UFFD handler. - **#4 — bundle.canonical_memory_manifest lifts onto SandboxSpec**: validates the wire shape so the canonical ref survives bundle → spec → snapshot manifest plumbing. - **#5 — content-addressed dedup across sessions**: same bytes chunked twice produce identical hashes (the property the cross-session dedup story relies on). - **#6 — Linux + nbd-module-gated NBD daemon spawn** (`#[ignore]`, validates the disk-side end-to-end on the dev VM): builds a disk manifest, attaches a real NBD daemon against `/dev/nbd0`, drops cleanly. Pure Rust, no FC, no KVM, no real kernel — runs in 0.42s on every CI lane (macOS + Linux). The Linux+KVM-gated tests for real FC + real NBD live in `crates/engram-sandbox-firecracker/ tests/snapshot_uffd.rs` + the planned `nbd_chunked_disk.rs` (task #34). 206 unit + integration tests pass workspace-wide: - engram-host-agent lib: 85 - engram-host-agent adr_0007_e2e: 5 - engram-sandbox-firecracker lib: 38 - engram-image-builder lib: 15 - engram-core lib: 25 - engram-coordinator lib: 43 - (the new test file adds 5 + 1 ignored) clippy + fmt clean on macOS; Linux check clean via dev VM.
nikhilunni
added a commit
that referenced
this pull request
May 14, 2026
The proxy + iptables landed in e27c331, but several gaps in split-mode wiring left the filter dark. This commit closes all of them so a Firecracker guest in mode=coordinator + mode=host now sees `getent hosts api.anthropic.com` succeed (allowed) and `getent hosts evil.exfil.com` NXDOMAIN (denied) — observable via the proxy log line `DNS denied; responding NXDOMAIN ... reason=NotInAllowList`. **1. host-agent never ran `FirecrackerBackend::host_startup`.** Iptables rules are installed there (inter-VM DROP, host-LAN filters, proxy REDIRECTs, default-deny FORWARD, MASQUERADE). The coordinator wired it for mode=all; the host-agent main.rs hadn't. Result: every iptables chain was empty in split mode, so VM egress flowed through the kernel's default policies with no filtering and no NAT. Now the host-agent calls `fc.host_startup()` right after building the backend (mode=host parity with mode=all). **2. `RemoteHostClient::guest_ip` returned `None` unconditionally.** `api/sessions.rs:537` only calls `notify_session_policy` when `guest_ip(sandbox_id)` returns `Some(_)`. With the coord-side stub returning `None`, the host's proxy registry never received the per-session policy — DNS queries hit it as `reason=UnknownGuest` and got NXDOMAIN'd. Wire surface gains `RequestKind::GuestIp { sandbox_id }` + `ResponseKind::GuestIp { ip: Option<String> }`; `RemoteHostClient::guest_ip` now sends the unary and unwraps the reply. **3. `FirecrackerBackend::guest_ip` always dialed agentd via vsock.** Even with #2 above, the first call races the agent's boot — the 2s timeout often expires before agentd is ready, returning `None`, which loses the only opportunity to register the policy (`notify_session_policy` is best-effort, not retried). Fast-path: the /30 allocator already knows the guest's IP statically (network + .2), so return that and cache it. Skips a vsock RTT and the boot race. The vsock path stays as the fallback for post-restore lookups where `net` is empty. **4. `NotifyKind::SessionEgressPolicy` was dropped on the host.** Phase 1's wire layer had `RemoteHostClient::notify_session_policy` send the policy as a Notify, but the host-side `handle_notify` match arm logged "received unexpected SessionEgressPolicy; ignoring" — the comment claimed it was a coord→host frame but the dispatcher treated it as host→coord echo. The arm now calls `backend.notify_session_policy(policy).await`, which routes through `PooledBackend` into `engram-egress-proxy::Registry:: register`. `handle_notify` now takes `backend: Arc<dyn HostClient>` alongside the existing `Option<Arc<dyn NotifyHandler>>` so the dispatch can reach both. **5. host-agent didn't plumb `--egress-proxy-port` into the FC config.** Symmetric with the coord-side wiring in `coord/main.rs:359`. Now set when `cli.egress_proxy_port > 0`. Pair with `ENGRAM_EGRESS_PROXY_PORT=9443` (set in `deploy/dev-split/run-host.sh`). **6. Stale TAP interfaces from earlier sessions accumulated.** Each prior session left a `tap-engr-XXX` DOWN with a route for the /30. The kernel happily kept multiple `10.200.0.0/30` routes; new sessions could pick a DOWN tap and silently drop traffic. Not a code fix here, but documented in the dev-split runbook: clean stale taps between restarts with `for t in $(ip -br link show | awk '/^tap-engr-/{print $1}'); do sudo ip link delete $t; done`. **`engram-sandbox-firecracker::net::host_startup_lines`** gains a second `dns_port: Option<u16>` parameter so the REDIRECT target follows whatever port the egress proxy actually bound. Default 5353 (avoids systemd-resolved on 127.0.0.53:53). The matching field `egress_dns_port: Option<u16>` lands on `FirecrackerConfig`. Two new net::tests assertions: DNS REDIRECT uses the configured port, no-proxy mode keeps the legacy ACCEPT. Drives the bogus-key Claude flow through to actual auth-error text once the harness retries past Anthropic's rate limit on 401-flood sessions. Workspace: 710/710 tests pass on Mac. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 15, 2026
Five new POST routes on the coord that any pod can serve, replacing
the host-initiated bits of the WS protocol. The host-agent doesn't
call them yet — that wiring lands with the gRPC server in the next
commit, and the actual cutover (delete WS dial, delete WS server)
happens after both transports are dark-side-by-side.
POST /api/hosts/register
Replaces NotifyKind::Hello + initial upsert_host. Body carries
host_id + hostname + host_addr (gRPC dial URL) + agent_version
+ wire_version + optional cloud_metadata. Persists the row with
host_addr; upsert_host's COALESCE prevents heartbeat-shaped
writes from clobbering it.
POST /api/hosts/:id/heartbeat
Replaces NotifyKind::Heartbeat + HeartbeatAck. Same body shape
as the WS path (HostCapacityReport + local_snapshots +
running_sandboxes + draining); same dispatch (reconcile,
update_state on host_registry, touch_host_heartbeat); returns
HeartbeatResponse { server_time, revoked_sessions }.
POST /api/hosts/:id/auth/resolve-registry
Replaces RequestKind::ResolveRegistryAuth + AuthRequestHandler.
Body: { registry_host }. Returns { creds: Option<{username,
password}> } via the existing PgAuthResolver — no new logic,
just an HTTP wrapper.
POST /sessions/:session_id/harness-events
Replaces NotifyKind::HarnessEvent forwarding. One POST per
event (the session_id is in the URL; sandbox_id + event + at
in the body). Calls state::emit_harness_event with the same
semantics; back-to-back duplicate harness_idle de-dup and the
monotonic session_events.idx are already in place.
POST /api/hosts/:id/idle-eviction-candidates
ADR 0011 follow-up #2. Body: list of { session_id, sandbox_id,
idle_since }. Each candidate runs through
idle_evictor::evict_idle_session (the canonical pipeline, kept
coord-side because it needs the metadata store + chunk store).
The pipeline's registry-guard short-circuit makes the call
idempotent across coord pods.
All five sit behind the existing bearer-token middleware. No new
state, no new dependencies — each handler is a thin shim over
existing logic. 719/719 tests pass; lint clean.
The GrpcHostPool warm-up call on register is stubbed with a TODO —
the pool isn't on AppState yet, and adding it now without a
consumer would just be unused state. Wires up in the cutover.
nikhilunni
added a commit
that referenced
this pull request
May 15, 2026
ADR 0013 final cutover. The bincode-over-WebSocket coord↔host
transport is gone. Replaced with:
- **coord → host**: HTTP/2 + gRPC via `GrpcHostPool` on the
coord, served by the host-agent's tonic `HostService` server.
Bundled `start_agent(id, agent, policy)` is one atomic RPC
that applies the egress policy to the host's proxy registry
BEFORE spawning the agent — no more frame-ordering invariant.
- **host → coord**: plain HTTP/JSON POSTs to five endpoints
(`/api/hosts/register`, `/api/hosts/:id/heartbeat`,
`/api/hosts/:id/auth/resolve-registry`,
`/sessions/:id/harness-events`,
`/api/hosts/:id/idle-eviction-candidates`) that any coord pod
can serve.
Architectural moves:
- `HostRegistry`'s in-memory entry no longer holds a
`ConnectedHost` or `admin_client`. The pool owns the gRPC
channel; the registry stores scheduler state + a trait object
pointing at the pool's `GrpcHostClient`. Dispatch via the trait
is unchanged from coord's perspective.
- Idle-eviction *driver* moves to the host-agent
(`engram_host_agent::idle_evictor`) since its local
`HarnessHub` is authoritative for "last harness activity." The
*pipeline* (`evict_idle_session`) stays in
`engram_coordinator::idle_evictor`, invoked by the new
HTTP endpoint. Idempotent across pods via the existing
registry-guard short-circuit.
- `apply_egress_policy` is a separate trait method (and gRPC RPC)
for the no-agent egress-policy path. Distinct semantics from
`start_agent` — not a transitional shim.
- ADR 0011 follow-ups #1 (HostAdminHandler folded into HostService),
#2 (idle_evictor driver moves to host), #3 (acquire/release_shell
via HostClient) all land here.
Deletions (no backwards-compat artifacts left):
- `engram-protocol/src/{client.rs, server.rs, codec.rs, scheduling.rs}`
- `engram-protocol/src/wire.rs` slimmed to just WireExecRequest +
WireReapStats + WIRE_VERSION constant (the bincode payload
types still used inside gRPC bytes fields)
- `engram-protocol/tests/loopback.rs`
- `engram-coordinator/tests/wire_integration.rs`
- `engram-host-agent/src/dialer.rs` (WS dial loop)
- `engram-host-agent/src/ws_auth.rs` (WsAuthResolver)
- `engram-coordinator/src/api/hosts.rs::{connect, handle_connection,
AuthRequestHandler}` + axum↔tungstenite bridges
- `engram-coordinator/src/idle_evictor.rs::{spawn, run_once,
idle_ttl_from_env, idle_hard_ttl_from_env, DEFAULT_*_TTL_SECS,
DEFAULT_POLL_INTERVAL}` — the polling driver
- `engram-protocol::scheduling` module (unused)
- `NotifyKind`, `RequestKind`, `ResponseKind`, `Frame`, `StreamItem`,
`RemoteError`, `TraceContext`, `HostRequestHandler`, the
bincode `WIRE_VERSION` invariant — all gone
- `tokio-tungstenite` dep dropped from `engram-host-agent` and
`engram-protocol` (still used in `engram-coordinator` for the
browser↔guest ttyd shell proxy, a separate WS use case)
Tests:
- 689/689 pass (was 719; the WS loopback + wire_integration tests
deleted with the code they tested)
- One test updated:
`admin_reap_materialize_dir_reports_host_without_admin_client`
renamed to `..._reports_host_not_in_grpc_pool` — same
graceful-skip contract, just keyed on the pool now.
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
Captures the six issues surfaced during the 2026-05-19/20 M1.16 rollout investigation so the next session has a clean checklist to work from. Each item has symptom + root cause + proposed fix shape, ordered by rough priority. 1. Idle-evict orphan snapshot dirs (caller-layer leak) — fix #1 in 752aea3 covered PooledBackend::snapshot; the *caller* sequence (snapshot → register → destroy) has no equivalent cleanup hygiene. Each retry leaks 4 GiB; ~12 min fills a 99 GB disk. 2. SnapshotId::new() per retry is the wrong default for retry-shaped paths — proper fix is sandbox-keyed scratch with atomic promote on success. Subsumes #1. 3. Stale templates rows referencing missing snapshot blobs — surfaced immediately by the new Ops Agent host-agent logs. The POST /api/enabled-images cascade needs blob-resolvability verification before inserting; coord-side sweeper for runtime prune. 4. No disk-pressure floor on idle-evict — defense-in-depth backstop even with #1 + #2 fixed. 5. Refill failures don't surface to coord — host-internal warm-pool refill loop is silent to coord until disk fills. Add heartbeat per-template refill-failure signal. 6. (Closed by 3b6aec3) host-agent logs missing from Cloud Logging — documented as closed-via-this-incident so the diagnostic capability's provenance is clear. Follow-up: teach the engrams-prod-ops logs-host.sh skill script to prefer Cloud Logging over SSH+journalctl. Net effect: session b511cf9b's "29.5 s warm activation" was actually cold (warm pool empty due to #3), and the disk-fill follow-on was #1+#2. Tackling these in a fresh session has a clean ADR-rooted starting point. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
Extends the Known Issues block with a "Pointers for the next session" subsection so a fresh session can pick up without re-deriving where to start. Per-issue: - entry-point file paths + grep anchors - proposed module additions (coord sweeper for #3, gauge for #4, heartbeat field for #5) - a SQL snippet for verifying stale templates against BlobStorage Plus: - suggested attack order (#3 first to unblock warm-path test, then #1+#2 as one commit, then #5, then #4 as defense) - ready-to-paste Cloud Logging queries for refill failures, snapshot dir leaks, eviction failures - production state snapshot at write time so a future reader can diff against current state (FC hosts pgs9+tf4j on 3b6aec3, demo image at warm-75babf7, harness still on pre-M1.12 b9dd2d1, 5+ known stale templates) - pointers to engrams-prod-ops skill scripts No code change — purely documentation hygiene for the handoff. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
Session bb3dd147 surfaced two additional issues for the post-M1.16 follow-up queue: #6: Shell tab architecturally broken from k8s coord. Coord opens ws://<guest_ip>:7681/ws directly, but 10.200.0.0/24 lives behind FC host TAPs — coord pod has no route. Worked in dev (single-VM); never worked from k8s. Failure mode: connect to ttyd: IO error: Connection timed out. M1.16's per-VM netns + bake-time networking gave the VM an eth0 but didn't add a path from coord. Fix: tunnel the shell through host-agent's gRPC channel via a new ProxyShell bidi-streaming RPC. #7: NBD manifest version conflict still surfacing despite fix #2 (commit de0abcb). Either retry exhausts, hits a different code path, or has a bug in the retry loop. Self-recovering so lower urgency. Verify via re-read of disk_daemon/backend.rs::flush and a unit test with a mock that returns VersionConflict once. Pointers section gains entry points for both, and the attack order interleaves #6 as a self-contained parallel shippable alongside #1/#2. Production state snapshot block updated: bb3dd147 added to the list of confirming sessions (alongside b511cf9b). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 20, 2026
The prod incident on `engrams-fc-xngk` (ADR 0014 issue #1/#2): when idle-eviction's post-snapshot pipeline (record_snapshot → destroy → mark Idle) failed, the host-side idle_evictor re-POSTed the candidate every ~30s, each retry minted a fresh SnapshotId, and each wrote a new ~4 GiB dir at `/var/lib/engram/sandboxes/snapshots/<new-uuid>/`. The 99 GB disk filled in ~13 min (25 dirs leaked). Adopt the scratch+commit pattern from the ADR with a smaller blast radius than full blob-key renaming: the host tracks the just-produced snapshot_id in a per-sandbox `inflight_snapshots: DashMap` and exposes two new RPCs. - `commit_snapshot(sandbox_id)` — pipeline fully succeeded; clear tracking. Artifacts (local dir + per-snapshot opaque blobs) are now PG-owned. Idempotent. - `abort_snapshot(sandbox_id)` — pipeline failed; rm -rf the local per-snapshot dir AND delete the per-snapshot opaque blobs (state.bin, sidecar.json, working_set.json). Content-addressed chunks stay (the existing `chunk_gc` reaps them). Idempotent. - `snapshot(sandbox_id)` on retry — if a prior in-flight snapshot for this sandbox is still tracked, abort it BEFORE producing the new one. Closes the leak class even when the coord pod crashes between snapshot and commit/abort (the next retry self-cleans). Trait surface: - `SandboxBackend` gains `commit_snapshot` + `abort_snapshot` with default no-op impls so Process/VZ-dev backends inherit unchanged. FC pooled overrides both. - `HostClient` gets the same pair (default Ok), wired through `HostRegistry`, `LocalHostClient`, and the gRPC client/server. - New `CommitSnapshot` + `AbortSnapshot` RPCs in host_service.proto. Coord-side `evict_idle_session` (`idle_evictor.rs`) now calls `abort_snapshot` on every post-snapshot failure path (record_snapshot, set_session_status, emit SnapshotTaken / Evicted / StatusChanged) and `commit_snapshot` on full success. Tests added: - `pooled_backend::tests::snapshot_lifecycle_tests::*` — 4 tests covering commit-keeps-artifacts, abort-removes-dir-and-blobs, abort-is-idempotent, and the retry overwrite case. - `idle_evictor::tests::evict_idle_session_aborts_snapshot_when_record_snapshot_fails` — spy HostClient counts abort calls; MiniMeta injects record_snapshot failure; assertion: abort=1, commit=0. - `idle_evictor::tests::evict_idle_session_commits_snapshot_on_full_success` — inverse: commit=1, abort=0. All 747 tests pass; `just check` clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 22, 2026
ADR 0015's M2 section was the only "Proposed" milestone with a fully landed implementation. Update the status header, the abstraction table, the M2 body, the phased rollout, and the "what isn't working" item #2 to reflect what shipped: - Seven-commit chain `b0c8eca` / `0d928e4` / `aa215f8` / `d69d140` / `1734cb5` / `ee3b5a8` / `42649dd`, with the verification block showing the new lifecycle observed end-to-end on the dev-vm (`pending -> created -> active -> completed`). - The M2 body rewrites the proposal-language ("Probably an enum with explicit transition_to methods") into the as-built shape: legality table, `transition_session` helper, create-path Created -> Active, host-loss two-stage transition through HostLost, retired retry loop, and the explicit non-decision on typestate (won't carry the type-level state across DB-backed request boundaries). - Phase 3 of the rollout is now M2-shipped; M3 and the larger refactors slide one phase later. README's session-lifecycle box gets the new arrows (Created in the cold path; HostLost in the host-loss path; explicit DELETE -> Completed). Surrounding prose adds the "Active is honest" guarantee that pre-M2 was a quiet aspiration. Older docs that mention `SessionStatus` (DESIGN.md's ADR-0005 retirement callout, known-issues.md's Phase-7 resolved note, history.md, the ADR-0002 / 0005 / 0014 archives) intentionally left as-is — they're historical records of what existed at the time of those decisions, not live references. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 25, 2026
…ejoins chunked-disk tracking Adds `e2e_resume_rejoins_chunked_disk_tracking` to e2e_stack.rs. Pins both ADR-0016 §"Phase B failure mode to close" symptoms that commit 5 fixed. Test flow (snapshot → evict → resume → assert): 1. Create session against the demo chunked-OCI image. 2. **Probe cow-state up front** — if null, the runner can't exercise the NBD path (Blacksmith CI without nbd.ko). Emit `::warning::` + return; the snapshot/evict/resume API wiring is still exercised below the probe but the regression assertions skip. Pattern matches commit 4b's flush-now e2e. 3. dd + sync 8 MiB into the chunked-disk-backed rootfs. 4. flush_now — pin "applied" before the snapshot so the snapshot's disk_manifest is known to be non-trivial. 5. POST /sessions/:id/snapshot → DELETE /sessions/:id/local → POST /sessions/:id/resume. Equivalent to a full idle-eviction + resume cycle minus the 30s idle wait. 6. **REGRESSION CHECK #1** (Symptom 1 of the ADR): cow-state on the resumed session must be Some(...). Pre-commit-5 this returned null because the resumed sandbox was never inserted into nbd_sandboxes. Post-commit-5 the NBD-attach branch enters the new sandbox_id into the map → diagnostic populates. 7. **REGRESSION CHECK #2** (Symptom 2 of the ADR): write a different file (`post-resume.bin`) + sync + flush_now. Outcome must be `applied` with manifest_version > 0. Pre-commit-5 the resumed sandbox had no chunked-disk dirty buffer at all (materialize-to-file fallback) and the second eviction-snapshot errored with `non-canonical jail layout`. Post-commit-5 the resumed sandbox is a fresh ChunkedDiskBackend rebased on the snapshot's disk_manifest; writes go through it, flush_sandbox returns Some. Driver gains three new HTTP helpers (`snapshot`, `evict_local`, `resume`) that wrap the corresponding REST endpoints. Same `#[ignore]` + `ENGRAM_E2E_COORD_URL` gate as the other e2e_stack tests, so the existing test-e2e-stack CI lane picks it up via `--run-ignored ignored-only` without ci.yml changes. 788/788 workspace tests pass under just check. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
May 25, 2026
…manifest_* Phase C's pin set needs to enumerate "the chunks every enabled image points at." The bake's ManifestRef lives in bundle.json, an OCI layer that the host's ImageCache pulls on prefetch. Coord has no ImageCache — but it does have the ManifestRef in hand at materialize-time (api/enabled_images.rs::materialize_disk_chunks parses bundle.json there to write chunks at the correct ref). Persist it on the row so the pin-set query is a pure PG SELECT, no OCI round-trip per sweep. Migration 0036_enabled_images_disk_manifest.sql adds nullable disk_manifest_id + disk_manifest_version with a both-or-neither CHECK plus a partial index on disk_manifest_id (skips harness-only rows where the bake produced just a manifest layer, no chunked-disk artifact). EnabledImage struct gains Option<ManifestRef>; row::enabled_image_from_row parses it; upsert_enabled_image binds it; SELECTs include it. materialize_disk_chunks now returns Option<ManifestRef> instead of (), and enable_image + refresh_enabled_image stamp the result on the row before upsert. Harness-only images get None and drop out of the pin set naturally via the partial index. New MetadataStore::list_enabled_image_disk_manifest_ids — pin-set source #1. Default Ok(vec![]); PG impl returns DISTINCT (id, version) tuples wrapped as ManifestRef. Renumbered the trait doc comments so sources #1/#2/#3 align with their physical order in the sweep. Existing chunked-disk images keep NULL refs until re-enabled. Rollout-gated: ENGRAM_CHUNK_GC_ENABLED=0 on first ship plus 1 week dry-run window absorbs the gap. No back-compat shim — per the active development principle. just check clean: workspace fmt + clippy + tests pass. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
5 tasks done
nikhilunni
added a commit
that referenced
this pull request
Jun 2, 2026
…C guest The post-roll prod dogfood of the ADR 0027 playwright bundle surfaced three FC-guest environment bugs the clean-container dev-vm spike couldn't (the guest is not a stock distro init). The engine itself — aux-drive attach, presence re-anchor, init-shim mount, agentd activation — works; the browser just couldn't render or record until these are fixed. All three validated live in a prod dev_vm session (full open → snapshot → screenshot → engram-share, and open → video-start → chapters → video-stop → engram-share, producing real PNG + WebM artifacts). 1. No /dev/shm. The minimal devtmpfs /dev carried no POSIX shared-memory mount, so chromium's renderer crashed (Page crashed → about:blank; every navigation timed out). engram-init now mounts a tmpfs /dev/shm. 2. Playwright force-injects --disable-dev-shm-usage as a default launch arg (config `args` are appended, so it can't be dropped by omission, only via ignoreDefaultArgs). It pushes the renderer's shared memory off /dev/shm into the system tmpdir. The bundle cli.config.json now sets ignoreDefaultArgs: ["--disable-dev-shm-usage"] so chromium uses the real /dev/shm from fix #1. 3. Non-standard /tmp (0755 owned by the session uid, not the conventional sticky 1777). Playwright's video-artifacts temp dir + chromium's tmpdir shm fallback both land under /tmp and fail for any other uid (a root `engram exec`, or a caps-dropped renderer) — surfacing as "no videos were recorded". engram-init now chmod 1777 /tmp. Fixes #1/#2 are required for any browser use; #3 is required for video. Re-baking the FC-host image (init-shim) + the playwright bundle (config) + re-enabling rolls them. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This was referenced Jun 4, 2026
This was referenced Jun 10, 2026
This was referenced Jun 28, 2026
nikhilunni
added a commit
that referenced
this pull request
Jul 3, 2026
#2) The UPDATE that flips a session row to `created` + binds `sandbox_id` had no status guard and ignored `rows_affected`, so it returned `Ok(())` even when the row was deleted (DeleteSession) or requeued (stale-pending scanner) while the restore RPC that precedes this call was in flight. That silently binds a live sandbox onto a gone/inconsistent row instead of hitting the existing `Err` teardown arm in session_boot.rs, which destroys the now-orphaned sandbox. Add `AND status = 'pending'` to the WHERE clause and return MetaError::NotFound on rows_affected() == 0, matching the convention used elsewhere in this file (e.g. delete_registry_credential).
nikhilunni
added a commit
that referenced
this pull request
Jul 3, 2026
…rd (finding #2 cont'd) Completes the previous commit: the WHERE clause guard alone silently swallowed a lost race (0 rows matched) as Ok(()) unless rows_affected() is actually checked. This was split out of the prior commit by mistake during hunk staging; closing the gap here.
nikhilunni
added a commit
that referenced
this pull request
Jul 3, 2026
…before the metric lookup PR #556 review findings #1, #2, #3. Finding #1 [CONFIRMED]: the prompt->run-start histogram mixed a coordinator-process `Utc::now()` with a Postgres `created_at` column — two different clocks. Under skew, that biases every sample or (when the coord clock lags) silently drops the good ones via the negative-duration guard. Fix: `MetadataStore::prompt_received_at` is renamed to `prompt_received_seconds_ago` and now returns the elapsed seconds computed PG-side (`EXTRACT(EPOCH FROM (NOW() - created_at))`), so both ends of the measurement share one clock. Finding #2 [CONFIRMED]: the receipt lookup ran between `append_session_event(run_started)` and `events.publish(...)`, inserting a synchronous PG round-trip in front of the live SSE frame the ADR-0052 held user-echo waits on to render. Fix: publish first, then do the best-effort metric join — a pure reorder, same metric semantics. Finding #3 [CONFIRMED, no code change demanded]: the receipt write in `send_prompt_core` lands before any request validation, so a rejected or retried `SendPrompt` can leave a permanent, possibly duplicate, receipt row. This is spec-inherited from issue #527's emit-before-resume design; documented in both `send_prompt_core` and `prompt_received_seconds_ago` so the `DESC LIMIT 1` tradeoff (anchors on the last retry, not the user-perceived first ask) isn't lost. Deduplication is left to a follow-up issue. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc
nikhilunni
added a commit
that referenced
this pull request
Jul 6, 2026
… overlapped boot legs + prompt-over-wire (#566) * feat(coord): migration 0077 — session.selected_skills + fleet_catalog_changed NOTIFY trigger Part of issue #535 (create-as-a-plan). Groundwork for two later commits: - sessions.selected_skills (TEXT[]) persists a create's dynamic-mount selection so the queue scanner's boot re-prepare can reconstruct it (ADR 0055 TODO(P1-D): queued creates currently boot with base skills only, since the queue row never carried the selection). - notify_fleet_catalog_changed() + a trigger on hosts scoped to current_bundles changes (guarded by IS DISTINCT FROM, so it stays quiet across the few-seconds heartbeat UPDATE and only fires on an actual host-roll stamp change) backs the coordinator's boot-bundle cache invalidation, added in a follow-up commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): per-enabled-image boot bundle cache (issue #535 a) Combines the issue's steps 2+3 into one commit (the NOTIFY listener's invalidation calls need the cache type to exist first; splitting them would leave a dead intermediate compile state). - New `boot_bundle` module: `BootBundleCache` read-through caches, per enabled image, the parsed manifest + fetched base-snapshot record + resolved memory/vcpu budgets (today re-derived on every create), plus the fleet's baked bundle-name→sha catalog (today re-scanned via `list_active_hosts` up to three times per create). Both are TTL'd (30s belt-and-braces against a dropped PgListener notification). - `engram-postgres`: `upsert_enabled_image` / `soft_delete_enabled_image` now fire `pg_notify('enabled_image_changed', image_uri)` inside their existing transaction (delivered iff it commits); `delete_enabled_image` fires it best-effort after, mirroring `org_secret_changed`. - `pg_listener`: subscribes `enabled_image_changed` (invalidates one cache entry) and `fleet_catalog_changed` (invalidates the whole catalog — see migration 0077's trigger, landed in the prior commit). - `prepare_from_grpc` / `prepare_from_row` / `prepare_inner` / `fleet_bundle_catalog` rewired onto the cache: the strict (non-soft-deleted) vs. tolerant (`_any`) split moves to the two call sites (the cache always fills via the tolerant view), and `boot_on_reserved_host`'s per-create `get_snapshot` is gone — the snapshot record now rides `BootInputs.base_snapshot` from the bundle. Net: a warm-cache create now does zero `toml::from_str` calls and zero extra `list_active_hosts` scans beyond the one placement still needs (`candidates_for` — deliberately kept, per the issue's "conscious divergence": placement needs a heartbeat-fresh host view). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): one-transaction session write-set (issue #535 b) Collapses `reserve_placement` + `enqueue_session_create` + the boot/ enqueue paths' separate satellite-write chains into a single `MetadataStore::reserve_and_persist_create` call whose Postgres impl commits the ENTIRE write-set — the row (placed or queued) plus every satellite (sealed secrets, capabilities, integration policy, harness, selected skills) — in ONE `FOR UPDATE` transaction, before any host RPC. - `engram-core`: new `SessionCreateWriteSet` / `CreateDisposition` + `MetadataStore::reserve_and_persist_create` (replaces `reserve_ placement` + `enqueue_session_create`) and `transition_session_created` (a slim `pending → created` + `sandbox_id` UPDATE, replacing `create_ session_created`'s INSERT-or-UPDATE upsert — the row is now guaranteed to already exist). `Session` gains `selected_skills: Vec<String>`, fixing the ADR 0055 TODO(P1-D) gap: a queued create's boot re-prepare can now reconstruct its dynamic-mount selection instead of silently dropping to base skills. - `engram-postgres`: the transactional impl (extends `reserve_placement`'s FOR-UPDATE body); `get_session`/`list_active_sessions`/ `list_queued_sessions_fifo` project the new column. - `engram-coordinator`: `boot_prepared` seals secrets (KEK, pure crypto — has no place inside the DB transaction) and serializes the policy BEFORE calling `reserve_and_persist_create`, then dispatches on `CreateDisposition` — `enqueue_create` as a separate function is gone, its Queued-disposition handling folds into `boot_prepared`. `boot_on_reserved_host` now does exactly ONE write of its own (`transition_session_created`, since the sandbox doesn't exist until the restore RPC returns) — the FK-ordering bug class (a satellite write racing the row's own insert; the ADR 0051 forge-token regression) is dead by construction, not by "call it after the row" convention. - Every other `MetadataStore` impl (9 test/mock fixtures across 6 crates) updated: the 2 that exercise the real create path (coordinator HTTP + gRPC integration tests) got honest in-memory equivalents; the rest mirror their pre-existing `unreachable!()`/`unimplemented!()` convention for unexercised trait surface. - New Postgres-level tests (`placement_reservation_live_pg.rs`) proving the write-set's atomicity: a single call commits every satellite together, and a forced mid-transaction failure (duplicate session_id) leaves NOTHING from that attempt — not even satellites that would have followed the failing statement. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): overlap the independent boot legs with the restore RPC (issue #535 c) `boot_on_reserved_host` no longer serializes the restore RPC behind the env/egress work (or vice versa) — the two are independent (neither touches the other's inputs) and now run concurrently via `tokio::join!`: - Restore leg: `restore_base_on_host` (the ~0.4-0.7s VM-side work). - Env/egress leg: the per-spawn forge/upload broker-token mint (`inject_harness_env`) + the integration policy's Plane-B injection resolution (`resolve_inject_entries`, which can round-trip an external mint-provider API for a mint-mode connector) — this is where that external round trip moves OFF the serial tail. Both only need `session_id`/`image_ref`/`integration_policy`, not the sandbox; the broker-token FK has been satisfiable since `reserve_and_persist_ create` committed the row, well before this function runs. `build_egress_policy` splits accordingly: `resolve_inject_entries` + `build_observe_entries` (sandbox-independent, now called from the overlapped leg) stay as-is; the renamed `assemble_egress_policy` is the remaining sandbox-dependent half (`guest_ip` + final assembly). Also parallelizes `resolve_policy_secrets`' per-secret `SecretStore` round trips (order-insensitive — no secret depends on another) via `futures::future::join_all`, replacing the one-at-a-time loop. `transition_session_created` + the removal of `boot_on_reserved_host`'s satellite writes already landed in the prior commit (they're the same underlying change as the one-transaction write-set — splitting them would have left a dead intermediate compile state), so this commit is scoped to the actual leg-overlap + secret-resolution parallelization. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord,harness): prompt over the wire (issue #535 d) The initial prompt no longer rides `ENGRAM_INITIAL_PROMPT` env — every prompt, first or follow-up, is now a harness-protocol `Prompt` frame: - `api/prompt.rs`: factored the echo-then-forward core out of `send_prompt_core` into `deliver_prompt(state, session_id, sandbox_id, prompt_id, text)` — the user-echo-first ordering (load-bearing for web rendering) and the self-healing `deliver_with_reattach` forward, shared by every caller. - `session_boot::boot_on_reserved_host`: after the Active flip, mints a server-side `prompt_id` and calls `deliver_prompt` for the initial prompt — replacing the synthetic `prompt_id: None` event. Delivery failure past the reattach budget is `BootError::Started` (terminal, consistent with a `start_agent` failure): a session that can't receive the prompt that created it is broken. - `resolve_harness` no longer takes an `initial_prompt` param or inserts `ENGRAM_INITIAL_PROMPT`; `git grep ENGRAM_INITIAL_PROMPT` now returns nothing. - `engram-harness-claude`: `run_engine` drops the `initial_prompt` param — the pending queue starts empty and the first prompt arrives via `HarnessCommand::Prompt` like every other one. Updated the 8 unit tests that seeded an initial prompt through the deleted parameter to instead send it via `cmd_tx` post-spawn (and to expect the leading `Idle` the engine now emits before any prompt arrives, since the env-seeded fast-path — "start turn 1 with no leading Idle" — no longer exists). - `engram-host-agent/tests/e2e_harness.rs`: `capture_sink` now also returns a command sender so `drive_harness` can push the initial prompt as a wire frame instead of an env var — a hand-rolled minimal stand-in for `HarnessHub` (this test drives `SandboxBackend` directly, no coordinator/hub in the loop). The queued path unifies for free: `SessionCreateWriteSet::queue_prompt` (landed in the write-set commit) is already the durable prompt, and `prepare_from_row` threads it into the identical `boot_on_reserved_host` delivery path — no separate queued-prompt spelling. Deviation: did not extend `e2e_stack.rs`'s create-with-prompt test (`e2e_claude_with_bogus_key_surfaces_anthropic_auth_error`) to assert `prompt_id` threads onto `run_started` — it's quarantined (#403, excluded from the gating e2e lane) and requires a live `ENGRAM_E2E_GRPC_ADDR` stack this environment doesn't have, so the change is unverifiable here. The same property (RunStarted.prompt_id matches the delivered prompt_id) is covered by engram-harness-claude's `queue_holds_edits_and_consumes_type_ahead` unit test instead. Note: engram-harness-claude's test module is `#[cfg(target_os = "linux")]`-gated and e2e_harness.rs is Linux+KVM+FC+Docker+sudo-gated — neither compiles or runs on this macOS dev machine; both are verified by careful reading + will run for real in CI's Linux/FC lanes. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * feat(coord): coord_prepare/coord_finalize phase metrics + doc pass (issue #535) Adds the two phase labels the issue's acceptance criteria need to turn "the coordinator serial tail is ~1s" from an estimate into a measurement, on the existing `engram_session_boot_seconds` histogram (no new metric names): - `coord_prepare`: `create_session_core` entry through `reserve_and_ persist_create`'s commit — the serial coordinator-side prefix ahead of the (now-concurrent, host-side) restore work. Recorded on the Placed path only. - `coord_finalize`: the restore RPC returning through the `created → active` transition — the coordinator-owned tail after the host hands back a live sandbox. Success path only. `total` minus (`coord_prepare` + `coord_finalize`) is the actual host-side restore RPC wall time — the split this issue's evidence section was missing. Doc pass: `session_boot.rs`'s module header now describes the (a)-(d) pipeline shape instead of the pre-refactor procedure; the FK-ordering guard comment in `prepare_inner` (the anchor the issue tracked as `sessions.rs:1197-1203`, drifted slightly by the time this landed) rewritten to describe the current dead-by-construction invariant instead of the historical hazard. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01T4WkZddtVw8djsWz2RCcQc * fix(coord): guard transition_session_created on status=pending (finding #2) The UPDATE that flips a session row to `created` + binds `sandbox_id` had no status guard and ignored `rows_affected`, so it returned `Ok(())` even when the row was deleted (DeleteSession) or requeued (stale-pending scanner) while the restore RPC that precedes this call was in flight. That silently binds a live sandbox onto a gone/inconsistent row instead of hitting the existing `Err` teardown arm in session_boot.rs, which destroys the now-orphaned sandbox. Add `AND status = 'pending'` to the WHERE clause and return MetaError::NotFound on rows_affected() == 0, matching the convention used elsewhere in this file (e.g. delete_registry_credential). * fix(coord): check rows_affected on the transition_session_created guard (finding #2 cont'd) Completes the previous commit: the WHERE clause guard alone silently swallowed a lost race (0 rows matched) as Ok(()) unless rows_affected() is actually checked. This was split out of the prior commit by mistake during hunk staging; closing the gap here. * docs(coord): fix stale FOR UPDATE lock-duration comment (finding #3) "Held only for the pick + insert below (sub-ms)" stopped being true once reserve_and_persist_create's satellite writes (sealed-secrets insert, per-capability insert loop, integration-policy upsert) moved inside the same transaction as the FOR UPDATE host-row lock (issue #535 (b)) — the lock is now held until tx.commit() at the end of the function, across all of that. Correct the comment so the next reader doesn't under-estimate placement-lock contention on a many-capability create burst. * refactor(coord): delete dead upsert_session_secrets/set_session_harness (finding #4) Both are zero-caller writers left behind by reserve_and_persist_create subsuming the old satellite-write paths: the sealed-secrets bytea and the harness selection now ride the one-transaction write-set / row INSERT directly (engram-postgres/src/lib.rs's reserve_and_persist_create), not these standalone upserts. Confirmed zero callers workspace-wide (including orchestrator/ and web/) before deleting — only the trait declarations, the PostgresStore impls, and 10 mock impls referenced them. Per the repo's clean-break convention, retire both from the trait, the PG impl, and every mock rather than leaving them as an orphaned, unused re-entry point for the FK-ordering/partial-write bug class this PR set out to kill. get_session_secrets/delete_session_secrets and get_session_harness are untouched — those remain live (read/delete) call sites. * fix(coord): split INSERT/UPDATE migration triggers to fix invalid WHEN-OLD DDL (finding #1) migration 0077's `hosts_notify_fleet_catalog_changed` trigger's WHEN clause referenced OLD on an AFTER INSERT OR UPDATE trigger. Postgres rejects this at CREATE TRIGGER time ("INSERT trigger's WHEN condition cannot reference OLD values") — OLD doesn't exist on INSERT and the restriction is static, not runtime, so the `OLD IS NULL` guard didn't help. This is the exact error CI hit ("while executing migration 77") and would crash-loop every coordinator replica at boot on merge, since migrations run at coordinator startup. Split into two triggers: an INSERT trigger with no WHEN clause (a new host's first bundle stamp always counts as a "change"), and an UPDATE trigger with `WHEN (OLD.current_bundles IS DISTINCT FROM NEW.current_bundles)`. Also updates the pg_listener.rs comment describing the guard now that it only applies to the UPDATE leg. * chore(migrations): renumber 0077 -> 0082 (batch land-queue collision) Six PRs in this land batch each added a migration numbered 0077. Land-queue assignment: #560 keeps 0077, #561->0078, #563->0079, #564->0080, #565->0081, this PR (#566)->0082. Pure rename plus updating the two in-repo comments that named the migration by number; no SQL content change. --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
… oracle, the gap-family seeds TTL clock -> now_mono: MigrationExport carries the injected clock; created_at/last_activity are now_mono readings and expired() subtracts against the injected clock — expiry DECIDES destroy/abort, so it is decision-feeding time (D1), off the metrics_now carve-out it previously rode. The prod constructors bind PooledBackend.clock; the Linux-only peer page server (PeerExport, which shares the same Arc anchor — #216 Gap 2) converts with it, carrying the same injected clock. The paused sim clock now drives the REAL expired() deterministically. (The musl cross-check caught the Linux-only PeerExport half — the macOS sweep can't see migrate_peer.rs.) Sim (engram-dst-host): steps MigrationBegin (a REAL MigrationExport in the REAL MigrationRegistry; deterministic entropy-minted export id — the prod OsRng nonce must not launder into the replayable id stream; the guest freezes exactly as the export's held capture lock excludes writes/flushes/captures — and the Flow F interleaving steps now gate on it, which also fixes a hang the swarm found: a flush step on a frozen slot armed a seam an empty pipeline never reached), MigrationServeState (the split-brain flag), MigrationTouch, MigrationTtlSweep (REAL expired() + ttl_verdict over the scriptable coordinator's ownership answer, applied as lib.rs does), MigrationCommit, MigrationAbort — in both swarm profiles (pick roll widened 0..112; old seeds re-explore, fine per the seed contract). Oracle #7 (the #216 decision table): state_served => never abort-unpause — a split-brain un-pause is structurally recorded by the sweep/abort appliers, so a ttl_verdict regression or bypassing caller fires it. Oracle #2 (no-plane-leak): migrating <=> an open registry export, with a live backend — no frozen guest ever leaks without an export to end it. Seeds: ttl_expired_unshipped_export_aborts_in_place_zero_loss, state_served_export_never_unpauses_then_destroys_on_ownership_flip, actively_serving_export_never_expires_mid_transfer (#216 Gap 1), unreachable_coordinator_stays_paused_never_guesses, reattached_source_verdict_never_destroys_on_a_transient_binding (Gap 3). Scope notes (recorded in the ADR row): #582/#598/#629 turned out to be FC-lane/test-hygiene issues whose portable content P5/P7 already absorbed — no hollow seeds manufactured. HostEffects::production consolidation deliberately retired rather than done: every seam reaches its flow through its own field; the bundle ctor remains the sim's assembly point (the pooled TODO now says so). ADR: P8 row -> Landed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Jul 18, 2026
Merged
nikhilunni
added a commit
that referenced
this pull request
Jul 18, 2026
…e model oracle (#786) * ADR 0098 R2: the host-effect queue (in-flight interruption + message faults) Pre-R2 every SimHostClient verb mutated world truth inline after one optional delay, so RPC loss/reorder/duplication and a replica crash BETWEEN the store commit (the ack the coordinator already holds) and the host-side world effect were structurally impossible (the audit's finding #2/#4). A mutating verb now records its world mutation as a typed `Effect`. Inline delivery is the default — byte-for-byte the pre-R2 world, so Calm seeds are unchanged (Calm never opens a deferred window). When the target host is in the deferred set (a Chaos fault window), the effect is queued under a monotonic serial (BTreeMap → deterministic order) for a later scheduler step to deliver / drop / duplicate / reorder: - DeferHost(i, on) — open/close a host's deferred window - DeliverEffects — deliver the queue in serial (causal) order - DropEffect — loss: drop one queued effect (seeded pick) - DuplicateEffect — re-deliver one effect (seeded pick) - ReorderEffects — deliver in a seeded-shuffled order Crash/restart of a host severs its in-flight effects (never resurrected onto the cleared VM set); quiescence closes every window and flushes the queue in order before the fleet heals, so world truth is consistent with the coordinator's commits. The Chaos weight table carves 9 points out of AdvanceTime/Driver/HostHeartbeats/Crash/RestartHost for the new arms, which shifts every Chaos seed's exploration (seeds pin to a commit); the pinned chaos seeds still converge and are kept. tests/effect_queue.rs pins the two reachable states the queue unlocks: a deferred create withheld-then-delivered, and a replica crash inside the commit→effect window followed by effect loss, both converging at quiescence (no_op_dropped + no-stragglers + no-orphans). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ADR 0098 R2: API-driven workload over the real Router/gRPC surface Pre-R2 the workload called store/core fns directly (0/17 HTTP, 0/~69 gRPC driven — the audit's finding). This lands the real surface: - workload.rs drives each replica's ACTUAL surface: the tonic AppSessionService/AppFleetService impls (auth + convert.rs + the same *_core the axum handler calls — handler-direct, no socket, to keep the paused-clock current-thread determinism) for create/prompt/resume/ delete/drain, and the actual axum api::router via tower::oneshot (real-wire, middleware + extractors) for admin pause/resume and the harness-idle host ingestion. A module honesty table documents which verb takes which path. - tests/api_surface.rs exercises it deterministically: create is acked, persisted (never lost), and REPLAYS byte-identically; prompt+delete round-trip; the axum router runs the real HTTP handlers to a normal response. Real bug found + fixed: create_session's prepare path minted the session id via a raw `SessionId::new()` (Uuid::new_v4) instead of the injected `services.entropy` — an ADR 0098 D1 determinism leak (the id diverged every replay; in prod OsEntropy makes this behavior-identical). The API-create path now replays. SimMeta gains four PG-faithful methods the real handlers reach (previously panic-stubs): insert/get/delete_broker_token, get/set_teleport_target, rebind_session_guarded, list_active_assignments_with_budgets_on_host — with broker-token and teleport-target conformance scenarios (ADR 0098 D4). Scope note: folding this workload into the chaos/calm SWARM additionally needs host-fidelity fixes (schedulable wire_version, staged ready_images) that unmask a genuine, separate driver-double-boot + eviction-durability class the pre-R2 sim silently suppressed by leaving hosts wire-skewed and snapshots non-recoverable. Those are real findings for a follow-up; here the surface is driven directly and deterministically rather than shipping a red swarm lane. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * ADR 0098 R2: the expected-state model oracle (the auditor) TigerBeetle's auditor shape: an expected-state model fed ONLY by ACKED workload outcomes, diffed against world/SimMeta truth in the standing invariant pass (every step + at quiescence). The honesty boundary is explicit — because the model is fed only by acks, a lost-response op (a create whose boot never durably established, an op whose reply dropped) is legitimately absent from the model and never asserted on. What it asserts: - acked-create/resume never silently lost: a session the workload saw reach Active (durable row committed) still has a row on every later step, unless a later ACKED destroy retired it. A row vanishing under a live session is exactly the #570 symptom class (coordinator-unbind vs host-teardown-reconcile racing a session out of existence) and the durability-lie class this program exists to catch. - read-your-acked-writes: the image a session was created with is never repainted (both replicas read the same shared SimMeta, so a present row is readable on either). Fed from the op-path CreateSession/ResumeSession outcomes (record_if_live records a session once observed Active). Wired into run()'s per-step and quiescence checks alongside the existing oracles. Non-vacuity is proven, not assumed: tests/model_oracle.rs drives a session to Active, drops its row directly (the corruption the #570 race produces), and asserts the auditor FIRES with `model-acked-create-not- lost`; restoring the row clears it. `SimMetadataStore::with_db_mut` is the test-only corruption injector. Swarm (chaos 0..40 x600) stays green — the auditor is a safety invariant the current code upholds; the canary test is what proves it can bite. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 21, 2026
…ession The test deliberately writes request #2 into the torn-down tunnel; which teardown error the proxy task surfaces is a platform/timing coin flip — macOS yields close_notify/UnexpectedEof, Linux CI yields Broken pipe. All mean the same thing the test asserts: the connection died. The substantive assertions (Connection: close rewrite, injected first request, no second request upstream, client EOF) are unchanged. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
nikhilunni
added a commit
that referenced
this pull request
Jul 21, 2026
…846 review) The Upgrade exemption in force_connection_close was itself a bypass: the header is guest-supplied, and any REST/GraphQL upstream that doesn't upgrade ignores it and keeps the connection persistent. A guest could dress gated request #1 with `Upgrade: websocket` + `Connection: Upgrade` and stream an ungated request #2 through copy_bidirectional — reopening the exact keep-alive gate bypass this PR closes (no method/path gate, no inject policy, no placeholder substitution on #2). Strip Upgrade like the other connection-management headers and always force `Connection: close`. No intercepted (credential/observe) host speaks websockets; genuine support would need a per-policy opt-in plus a 101-aware tunnel, not trust in a client header. New e2e regression drives the actual attack (Upgrade-dressed #1 → dead #2); unit test pins the strip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bumps the docker-base group with 1 update in the /docker directory: rust.
Updates
rustfrom 1.83-slim to 1.95-slim