Repository navigation
feat(uffd-handler): post-copy demand faults fetch a window of chunks (ADR 0045 C2 addendum) - #1611
Open
nikhilunni wants to merge 1 commit into
Open
nikhilunni wants to merge 1 commit into
nikhilunni wants to merge 1 commit into
Conversation
…(ADR 0045 C2 addendum) A live move of an 8 GiB guest reached Active only after a 22 s harness handshake: every page agentd touched was one remote fault, 12,727 of them at 8 ms each. A demand fault now sends one `NeedWindow` for the faulting chunk plus up to seven sealed, uninstalled neighbours in its region (`--fault-around-chunks`, env `ENGRAM_UFFD_FAULT_AROUND`, default 8, range 1..=256; 1 is the old behaviour). - The fault loop installs and wakes only the faulting chunk from the first streamed response, then returns. A prefetch worker on its own fault connection installs the suffix through the existing install claims, outside the fault-active guard, so a slow suffix never stalls another vCPU or the drain. An overlapping demand goes out as a one-chunk window on the demand connection and the claims deduplicate. - Every response shape is validated against the request's `req_id` and expected `chunk_offset`; a reordered or duplicate reply marks the peer lost before any install. Retries request only the unfinished suffix. - Worker connects are bounded and cancellation-aware; `stop` covers lock acquisition and the join under one deadline and detaches with a warn; blocking fetches run outside the install lock with a cancellation recheck before any copy. - `fault_around_chunks_installed` rides the handler stats, the control report and the host's fault-path totals. Page-channel `PROTO_VERSION` 5 -> 6; coordinator `WIRE_VERSION` 34 -> 35 so hosts roll in lockstep. Tests: streamed installation, suffix-only retry, corrupt and reordered replies, one-chunk equivalence with `NeedAt`, window selection over random seals (proptest), concurrent drain dedup, faulting-page-only wake, overlapping and unrelated faults and the drain bypassing a stalled suffix, cancellation during a pending connect, a worker that misses the stop deadline, a stalled AltSource fetch under stop; the two-host KVM test checks the counter is reported. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM
nikhilunni
force-pushed
the
feat/postcopy-fault-around
branch
from
October 8, 2026 00:03
081008e to
2174654
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Post-copy demand faults now fetch a window of chunks in one round trip (ADR 0045 C2 addendum). Measured in production: a live move of an 8 GiB guest reached Active only after a 22 s harness handshake, because every page agentd touched was a separate remote fault (12,727 faults at 8 ms each). This amortises those round trips.
The change
NeedWindow: the faulting chunk first, then up to seven sealed, uninstalled neighbours in the same region.--fault-around-chunks/ENGRAM_UFFD_FAULT_AROUND, default 8, range 1 to 256; 1 reproduces the old single-chunk behaviour.req_idand expectedchunk_offset; reordered or duplicate replies mark the peer lost before any install. Retries request only the unfinished suffix.stopcovers lock acquisition and the join under one deadline and detaches with a warning. Blocking fetches run outside the install lock with a cancellation recheck before any copy.fault_around_chunks_installedis reported in the handler stats, the control report and the host's fault-path totals.PROTO_VERSION5 → 6. CoordinatorWIRE_VERSION34 → 35 so hosts roll in lockstep, as the previous page-channel bump did.Drain order, hot-first logic, seal rules and DrainDone/PeerLost handling are unchanged.
Tests
Streamed installation; suffix-only retry; corrupt, reordered and duplicate replies; one-chunk equivalence with
NeedAt; window selection over random seals (proptest, pinned persistence); concurrent drain dedup; faulting-page-only wake; overlapping and unrelated faults plus the drain bypassing a stalled suffix; cancellation during a pending connect; a worker that misses the stop deadline; a stalled AltSource fetch under stop. The two-host KVM test checks the counter is reported.Validation
just checkgreen (2771 tests); Linux cross clippy for uffd-handler, host-agent, migrate-proto and protocol clean; four Codex adversarial review rounds, final verdict approve.Follow-up
The next prod measurement should show the handshake time on a live move drop from 22 s. The idle-page hot set PR (separate) gives the drain a recency order for File-mode sources so fewer faults happen at all.
🤖 Generated with Claude Code
https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM