Repository navigation
chore(web)(deps-dev): bump typescript from 5.9.3 to 6.0.3 in /web - #7
dependabot[bot] wants to merge 1 commit into
Conversation
|
@dependabot recreate |
f3b5486 to
3c87b08
Compare
|
@dependabot rebase |
3c87b08 to
94f7b6c
Compare
|
@claude please fix the failing |
The idle evictor only fires in --mode=all because the coordinator's HarnessHub stays empty in multi-host mode (RemoteSandboxBackend's set_harness_sink is a trait default no-op). For v1 production deploys, manual flush via /api/admin/sessions/:id/flush and /api/admin/flush-idle covers the gap; a CronJob can run flush-idle periodically. The proper v2 fix (instantiate HarnessHub on the host-agent, wire session binding, ship candidates over WS via a new NotifyKind) is the same cross-machine state-sync work that broker mode needs. Captured in known-issues #7 and deploy.md. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…reemption) #7 captures the bug pattern observed today — Active sessions sticking to dead sandbox_ids after a coord/host restart — and forward-references ADR 0009's Phase 3 reconcile pass as the fix. Retire once Phase 3 lands. #17 captures the ephemeral-host preemption gap (spot preemption, autoscale-down) that ADR 0009 explicitly defers. Notes that preemption-time snapshot is newly feasible in the chunked-memory world and forward-references the future preemption-drain ADR. Renumbers #8–#16 accordingly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sweeps every doc that the recent coord↔host work made stale and adds the missing pieces. No code changes. **Updates:** - `README.md` — `WIRE_VERSION=5`/`4` → `1`; the `RequestKind`/ `ResponseKind`/`NotifyKind` lists now reflect the post-Phase-1 shape (`StartAgent`, `BindHarnessSession`, `UnbindHarnessSession`, `SendHarnessPrompt`, `GuestIp`, `HarnessOk`, `HarnessEvent`); the host-agent description now frames `HostClient` as the coord↔host trait and `SandboxBackend` as the local VMM driver. - `DESIGN.md` — same trait clarification on the host-agent section, on the PooledBackend description, and the trait-design block — which now carries a `HostClient` section ahead of the (narrowed) `SandboxBackend` section. - `docs/demo-firecracker.md` — the iptables block had a stale `ACCEPT VM→DNS to 1.1.1.1` line and no mention of the DNS REDIRECT or the proxy's UDP/TCP-53 listeners. Rewritten to match what `engram-sandbox-firecracker::net::host_startup_lines` installs today. Bottom of file now cross-links to the new split-mode runbook. - `docs/deploy.md` — env-var table for the host-agent gains `ENGRAM_EGRESS_PROXY_PORT` + `ENGRAM_EGRESS_CA_SOURCE`. Topology section names the mode flags; egress section gains a "DNS filtering" subsection that points at ADR 0010 and the DNS-over-HTTPS residual gap. The "idle auto-eviction is a v2 feature" warning updated to reflect that the routing path is in place after ADR 0011 — the driver-shape fix is the remaining piece. - `docs/known-issues.md`: - #7 (active sessions stick after restart) → FIXED, ADR 0009 phases 1-8 landed. - #8 (idle eviction in `--mode=coordinator`) → reframed: the `HarnessHub` routing landed with ADR 0011; the remaining gap is the driver shape, with the recommended fix flipped to "move driver to host-agent" (the eviction TTL clock should be authoritative on the host). - #15 (wire compatibility): `WIRE_VERSION` is `1` after the not-yet-deployed reset; the historical changelog is gone. **New:** - `docs/adr/0010-dns-filtering.md` — DNS proxy on udp/5353 + tcp/5353, iptables REDIRECT, port choice (avoiding systemd-resolved), the DNS-over-HTTPS residual hole that can't be closed at the transport layer, the rejected alternatives. Extends ADR 0006. - `docs/adr/0011-host-client-trait.md` — the trait split: why `SandboxBackend` was carrying two abstractions, the five asymmetries that surfaced in split-mode, the resulting trait shape (`HostClient` = coord↔host boundary; `SandboxBackend` = local VMM driver), the wire-surface additions (`StartAgent`, bind/unbind/send_prompt, `GuestIp`, `HarnessEvent`), and the out-of-scope follow-ups (idle_evictor, shell.rs, HarnessHub location). - `docs/demo-split-mode.md` — runbook for `--mode=coordinator + --mode=host` on the dev-vm. Covers the launch scripts in `deploy/dev-split/`, the 10-step request flow that exercises every Phase 1-4 commit, the allowed-vs-denied DNS test, and the stale-TAP cleanup between restarts. Cross-linked from `docs/demo-firecracker.md`. - `docs/adr/0006` gains a "Follow-up: DNS filtering (ADR 0010)" pointer. - `docs/adr/0009` status bumped: "proposed (review pending)" → "accepted (phases 1-8 landed)". Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Session bb3dd147 surfaced two additional issues for the post-M1.16 follow-up queue: #6: Shell tab architecturally broken from k8s coord. Coord opens ws://<guest_ip>:7681/ws directly, but 10.200.0.0/24 lives behind FC host TAPs — coord pod has no route. Worked in dev (single-VM); never worked from k8s. Failure mode: connect to ttyd: IO error: Connection timed out. M1.16's per-VM netns + bake-time networking gave the VM an eth0 but didn't add a path from coord. Fix: tunnel the shell through host-agent's gRPC channel via a new ProxyShell bidi-streaming RPC. #7: NBD manifest version conflict still surfacing despite fix #2 (commit de0abcb). Either retry exhausts, hits a different code path, or has a bug in the retry loop. Self-recovering so lower urgency. Verify via re-read of disk_daemon/backend.rs::flush and a unit test with a mock that returns VersionConflict once. Pointers section gains entry points for both, and the attack order interleaves #6 as a self-contained parallel shippable alongside #1/#2. Production state snapshot block updated: bb3dd147 added to the list of confirming sessions (alongside b511cf9b). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ADR 0014 issue #7. The retry guard added in `de0abcb` for the NBD manifest version race used `if attempted - latest > 32` over `u64` operands. The typical race shape is the store being AHEAD of us (another sandbox already published version N+1 between our snapshot and our flush), so `attempted < latest` and the subtraction wraps to ~u64::MAX. In debug builds the wrapping_sub panics; in release builds it returns a huge value, the `> 32` check trips, and the very first retry bails out — returning the original VersionConflict instead of re-targeting `latest + 1` and re-PUTing. That's why `nbd disk flush: chunk store: manifest <id> version conflict: attempted v2, latest is v3` continued to surface in post-M1.16 prod logs despite the prior retry fix. Replace the directional subtraction with an explicit retry counter (`MAX_FLUSH_RETRIES = 32`). Same intent ("give up after a reasonable number of attempts") expressed in a way that doesn't depend on which side of the conflict is ahead. Tests added: - `flush_retries_past_version_conflict_with_store_ahead` — direct regression for the bug. Pre-publishes v2, dirties bytes, flushes; must observe `latest=2` and retarget v3. Pre-fix the same test would have bailed immediately on attempt v2 due to the u64 wrap. - `flush_retries_past_multiple_foreign_versions` — multi-level race where two foreign versions get pre-published, retry observes latest jumping forward and converges on v4. All 749 tests pass; `just check` clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
P4 was framed as 'pre-allocate a netns pool (E2B network/pool.go shape)'. Reading E2B's actual implementation (packages/orchestrator/pkg/sandbox/network) shows their speed comes from three compounding choices we lack — the pool is their endgame, not the first lever: - the ~91ms netns leg is mostly ip/iptables subprocess fork+exec; E2B uses netlink (vishvananda/netlink) — microseconds each, no subprocess - constant in-namespace identity (tap0 / 169.254.0.x) makes slots generic and reusable; we bake a per-template TAP name + /30 into the snapshot - per-sandbox config is firewall-only (nftables CIDR rules), often a literal no-op Reframed ladder: (a) subprocess->netlink swap first (biggest cheap win, zero reuse-correctness risk), (b) constant in-ns identity (bake-side, the reuse prereq), (c) pool last. Parked: ~40-60ms yield on a path the warm substrate (~516ms) + agent first-token gap still dominate. Also park #7 (overlap prefetch with restore_in_jail): doubly superseded — memory prefetch is already cheap (P2 prod finding) and ADR 0021 residency retires the boot-path prefetch outright. Updated the live phase-status tracker. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Bumps [typescript](https://github.com/microsoft/TypeScript) from 5.9.3 to 6.0.3. - [Release notes](https://github.com/microsoft/TypeScript/releases) - [Commits](microsoft/TypeScript@v5.9.3...v6.0.3) --- updated-dependencies: - dependency-name: typescript dependency-version: 6.0.3 dependency-type: direct:development update-type: version-update:semver-major ... Signed-off-by: dependabot[bot] <support@github.com>
94f7b6c to
dd85a44
Compare
… the lifecycle kernel (#543) (#599) * feat(ops): ADR 0079 — durable per-session op log with fencing epochs: the lifecycle kernel (#543) Every session lifecycle verb (resume, evict, deliver, create_boot, destroy) is now a durable session_ops PG row driven by a single-writer- per-session executor. sessions.current_epoch is the fencing epoch: CAS-bumped in the same transaction that claims an op, appended (AND current_epoch = $e) to every session-row write an op makes, and carried (SessionFence) on every session-scoped host RPC, gated by the host's persisted per-session high-water (work_dir/epochs, rejecting stale with FAILED_PRECONDITION). WIRE_VERSION 11 -> 12 (lockstep roll). The executor (session_ops.rs) is LISTEN/NOTIFY-hot (pg_notify 'session_ops'; 5s poll = fallback only), claims inline in the enqueue transaction on the idle-session happy path (one PG round trip), records durable per-step markers (idempotent-from-step crash resume), and replaces the lease reaper with fence-then-resume reclaim. Manual snapshot / evac resume / live teleport ride inline OpClaims on the same primitive pending their own verb phases. DELETED (grep-zero): the SessionLeaseGuard ecosystem + heartbeat + reaper + LeaseTouch taxonomy, the session_lease table (migration 0093) and its MetadataStore surface, both 8x3s lease-acquire retry loops, the Evicting hold + its polls/env, the resume/evict/snapshot spawn-detach pipelines, the residual-sandbox destroy compensation, the queue scanner's requeue-by-poll, and the prompt path's mid-move HOLD. The evict-then-resume collision class (3-21s stalls) is structurally gone: a resume behind an in-flight evict is ordering by log, zero sleeps. The deliver verb absorbs the outbox driver's per-session single-flight (rows/ack/202 framing stay ADR 0073's) and performs the ADR 0074 rung-1/2 ascent inline under its own fence. Epoch-0 reject is DEFERRED (allow-but-don't-advance) for the four named out-of-op senders — see the ADR divergence log and check_session_epoch's disposition table. Migrations 0092 (session_ops + current_epoch) + 0093 (drop session_lease). Tests: session_ops_live_pg (claim CAS, one-running, fencing 0-rows, ordering zero-sleeps, idempotency, cancellation, backoff-yields-head), session_epochs host tests, verb-level ordering units, migrated scanner/evicting-gate/outbox suites. Workspace 1501 passed; CI live-PG lane 106 passed; clippy/fmt/hakari clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014jJi2vqAaxt3Q5UKxbe4Gx * fix(ops): fence the evict/manual-snapshot record so a reclaimed-out op can't land a phantom recoverable row (ADR 0079 re-review #3/#4) The `park_or_capture` eviction step brackets pause → snapshot → record → commit with no intermediate `ctx.step()` fence check, and the capture leg can run minutes. A coord↔PG partition that outlasts `RECLAIM_STALE` (180s) lets a successor op re-claim the session (CAS-bumping `current_epoch`) while the predecessor is mid-capture. `record_snapshot` was a PLAIN, unfenced INSERT: on partition-heal the fenced-out predecessor would land a `recoverable` row a resume could pick (the 89f7984d durability-lie class) and then issue `commit_snapshot` under the stale epoch — the host's per-session epoch high-water is only eventually-consistent with PG (a reclaim bumps PG but not the host until the successor's first fenced RPC), so that commit can slip through. Add `MetadataStore::fenced_record_snapshot(snap, epoch)`: it writes the row ONLY while `sessions.current_epoch == epoch`, atomically in one transaction (the fence read is `FOR UPDATE`, serializing against the claim/reclaim CAS), returning Ok(false) when fenced. The Postgres impl routes both `record_snapshot` and the fenced variant through one private `record_snapshot_guarded(snap, fence: Option<i64>)` so the INSERT + generation bump + durable-head advance stay byte-identical; the default trait impl delegates (fence-less) for the in-memory mocks. The eviction pipeline and the manual-snapshot `recoverable=true` promote now use it and bail cleanly (no phantom row, no `commit_snapshot`) when fenced. PG is the authority, so this closes the window WITHOUT the proactive host-epoch-advance-at-reclaim (ADR 0079 deferral #1c stays a pure optimization). Live-PG regression: `fenced_record_snapshot_writes_only_under_current_epoch`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * fix(migration): heartbeat the teleport claim's synchronous body + fix stale RECLAIM_STALE comments (ADR 0079 re-review #1/#5) `migrate_session_live` held its `OpClaim` across the entire synchronous move body (presetup → concurrent restore-await → blackout → rebind → reactivate) with NO liveness beat — the only `heartbeat_at` stamp was `try_acquire`'s, and the first refresh (`claim.touch`) is in the finalize drain loop, which runs AFTER the body. A move whose body outlives `RECLAIM_STALE` (180s: a large VM over a slow inter-host link) was therefore reclaimed out from under a PERFECTLY HEALTHY holder — no partition required — and that holder then kept driving its (un-fenced, #7) migration RPCs while the successor's reclaim failed the row: a double-drive. The manual-snapshot inline claim already beats its body via `spawn_heartbeat`; the teleport body now does too, dropped right before the finalize task takes over its own beat. Also fix two stale "60s staleness" comments (`RECLAIM_STALE` was raised 60s → 180s in review finding #1): the drain-loop touch comment in live_migration.rs and the `OpClaim::spawn_heartbeat` doc in session_ops.rs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * fix(migration): fence the source-mutating migration RPCs so a stale teleport holder can't blackout/commit a live source (ADR 0079 re-review #7) `migration_capture`/`_presetup`/`_capture_postcopy`/`_commit`/`_abort` carried NO fencing_epoch, so `check_session_epoch` never gated them — the one class of session-scoped host RPC the host could not reject. Combined with the teleport claim's (now-fixed) missing heartbeat, a coord↔PG partition (coord↔host is a SEPARATE gRPC transport, ADR 0013) that outlasts RECLAIM_STALE lets a successor pod reclaim the teleport op while the original holder stays alive and drives migration_capture_postcopy (blackout) or migration_commit (destroys the source VM) with nothing able to reject it. Thread a SessionFence through all five: MigrationCapture/MigrationPresetup now take FencedSandboxRequest; MigrationExportRef gains fencing_epoch + session_id (MigrationCapturePostCopy/Commit/Abort). The host-agent gRPC server runs check_session_epoch on each and forwards the fence to the backend seam; the coordinator's teleport pipeline stamps claim.fence() at every call site. MigrationFetch (host-to-host, export-nonce gated) and MigrationDrainWait (dest-side) are deliberately unfenced — no source-mutating write on the holder's behalf. SandboxBackend stays fence-free (backend-local); only the HostClient wire seam carries it. No WIRE_VERSION bump — the fields ride the unreleased v12 (clean break). Also threads the fence through two Linux-gated FC teleport tests whose `restore`/`migration_*` client calls the base ADR 0079 commit left un-updated (macOS clippy can't see cfg(target_os="linux") — a green-local/red-CI hole); they pass SessionFence::unfenced() (epoch 0 is the host's allow-but-don't-advance interim). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K * docs(adr-0079): record the post-merge adversarial re-review fixes Adds a divergence-log section documenting the re-review of the two deferrals: the shared "a reclaim only fires on a genuinely dead executor" justification is false under a coord↔PG partition (ADR 0013 makes coord→host a separate transport). re-#1 (teleport body heartbeat), re-#7 (migration-RPC fencing, was deferred), and re-#3/#4 (fenced record_snapshot) are FIXED; re-#1c (proactive host-epoch-advance) stays deferred but re-grounded as a pure optimization now that PG-side fenced writes make the host high-water's eventual-consistency window harmless-by-construction; re-#5 (stale comments) fixed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BxMCSi2aBAJbHSMz9RNA7K --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
|
OK, I won't notify you again about this release, but will get in touch when a new version is available. If you'd rather skip all updates until the next major or minor version, let me know by commenting If you change your mind, just re-open this PR and I'll resolve any conflicts on it. |
… oracle, the gap-family seeds TTL clock -> now_mono: MigrationExport carries the injected clock; created_at/last_activity are now_mono readings and expired() subtracts against the injected clock — expiry DECIDES destroy/abort, so it is decision-feeding time (D1), off the metrics_now carve-out it previously rode. The prod constructors bind PooledBackend.clock; the Linux-only peer page server (PeerExport, which shares the same Arc anchor — #216 Gap 2) converts with it, carrying the same injected clock. The paused sim clock now drives the REAL expired() deterministically. (The musl cross-check caught the Linux-only PeerExport half — the macOS sweep can't see migrate_peer.rs.) Sim (engram-dst-host): steps MigrationBegin (a REAL MigrationExport in the REAL MigrationRegistry; deterministic entropy-minted export id — the prod OsRng nonce must not launder into the replayable id stream; the guest freezes exactly as the export's held capture lock excludes writes/flushes/captures — and the Flow F interleaving steps now gate on it, which also fixes a hang the swarm found: a flush step on a frozen slot armed a seam an empty pipeline never reached), MigrationServeState (the split-brain flag), MigrationTouch, MigrationTtlSweep (REAL expired() + ttl_verdict over the scriptable coordinator's ownership answer, applied as lib.rs does), MigrationCommit, MigrationAbort — in both swarm profiles (pick roll widened 0..112; old seeds re-explore, fine per the seed contract). Oracle #7 (the #216 decision table): state_served => never abort-unpause — a split-brain un-pause is structurally recorded by the sweep/abort appliers, so a ttl_verdict regression or bypassing caller fires it. Oracle #2 (no-plane-leak): migrating <=> an open registry export, with a live backend — no frozen guest ever leaks without an export to end it. Seeds: ttl_expired_unshipped_export_aborts_in_place_zero_loss, state_served_export_never_unpauses_then_destroys_on_ownership_flip, actively_serving_export_never_expires_mid_transfer (#216 Gap 1), unreachable_coordinator_stays_paused_never_guesses, reattached_source_verdict_never_destroys_on_a_transient_binding (Gap 3). Scope notes (recorded in the ADR row): #582/#598/#629 turned out to be FC-lane/test-hygiene issues whose portable content P5/P7 already absorbed — no hollow seeds manufactured. HostEffects::production consolidation deliberately retired rather than done: every seam reaches its flow through its own field; the bundle ctor remains the sim's assembly point (the pooled TODO now says so). ADR: P8 row -> Landed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Bumps typescript from 5.9.3 to 6.0.3.
Release notes
Sourced from typescript's releases.
Commits
050880cBump version to 6.0.3 and LKGeeae9dd🤖 Pick PR #63401 (Also check package name validity in...) into release-6.0 (#...ad1c695🤖 Pick PR #63368 (Harden ATA package name filtering) into release-6.0 (#63372)0725fb4🤖 Pick PR #63310 (Mark class property initializers as...) into release-6.0 (#...607a22aBump version to 6.0.2 and LKG9e72ab7🤖 Pick PR #63239 (Fix missing lib files in reused pro...) into release-6.0 (#...35ff23d🤖 Pick PR #63163 (Port anyFunctionType subtype fix an...) into release-6.0 (#...e175b69Bump version to 6.0.1-rc and LKGaf4caacUpdate LKG8efd7e8Merge remote-tracking branch 'origin/main' into release-6.0