Skip to content

feat(uffd-handler): post-copy demand faults fetch a window of chunks (ADR 0045 C2 addendum) - #1611

Open
nikhilunni wants to merge 1 commit into
mainfrom
feat/postcopy-fault-around
Open

nikhilunni wants to merge 1 commit into
mainfrom
feat/postcopy-fault-around

Conversation

@nikhilunni

Copy link
Copy Markdown
Contributor

Summary

Post-copy demand faults now fetch a window of chunks in one round trip (ADR 0045 C2 addendum). Measured in production: a live move of an 8 GiB guest reached Active only after a 22 s harness handshake, because every page agentd touched was a separate remote fault (12,727 faults at 8 ms each). This amortises those round trips.

The change

  • A demand fault sends one NeedWindow: the faulting chunk first, then up to seven sealed, uninstalled neighbours in the same region. --fault-around-chunks / ENGRAM_UFFD_FAULT_AROUND, default 8, range 1 to 256; 1 reproduces the old single-chunk behaviour.
  • The fault loop installs and wakes only the faulting chunk from the first streamed response, then returns. A prefetch worker on its own fault connection installs the suffix through the existing install claims, outside the fault-active guard, so a slow suffix cannot stall another vCPU or the drain. An overlapping demand is served as a one-chunk window on the demand connection; the claims deduplicate.
  • Every response shape is validated against the request's req_id and expected chunk_offset; reordered or duplicate replies mark the peer lost before any install. Retries request only the unfinished suffix.
  • Worker connects are bounded and cancellation-aware. stop covers lock acquisition and the join under one deadline and detaches with a warning. Blocking fetches run outside the install lock with a cancellation recheck before any copy.
  • fault_around_chunks_installed is reported in the handler stats, the control report and the host's fault-path totals.
  • Page-channel PROTO_VERSION 5 → 6. Coordinator WIRE_VERSION 34 → 35 so hosts roll in lockstep, as the previous page-channel bump did.

Drain order, hot-first logic, seal rules and DrainDone/PeerLost handling are unchanged.

Tests

Streamed installation; suffix-only retry; corrupt, reordered and duplicate replies; one-chunk equivalence with NeedAt; window selection over random seals (proptest, pinned persistence); concurrent drain dedup; faulting-page-only wake; overlapping and unrelated faults plus the drain bypassing a stalled suffix; cancellation during a pending connect; a worker that misses the stop deadline; a stalled AltSource fetch under stop. The two-host KVM test checks the counter is reported.

Validation

just check green (2771 tests); Linux cross clippy for uffd-handler, host-agent, migrate-proto and protocol clean; four Codex adversarial review rounds, final verdict approve.

Follow-up

The next prod measurement should show the handshake time on a live move drop from 22 s. The idle-page hot set PR (separate) gives the drain a recency order for File-mode sources so fewer faults happen at all.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM

…(ADR 0045 C2 addendum)

A live move of an 8 GiB guest reached Active only after a 22 s harness
handshake: every page agentd touched was one remote fault, 12,727 of
them at 8 ms each. A demand fault now sends one `NeedWindow` for the
faulting chunk plus up to seven sealed, uninstalled neighbours in its
region (`--fault-around-chunks`, env `ENGRAM_UFFD_FAULT_AROUND`,
default 8, range 1..=256; 1 is the old behaviour).

- The fault loop installs and wakes only the faulting chunk from the
  first streamed response, then returns. A prefetch worker on its own
  fault connection installs the suffix through the existing install
  claims, outside the fault-active guard, so a slow suffix never stalls
  another vCPU or the drain. An overlapping demand goes out as a
  one-chunk window on the demand connection and the claims deduplicate.
- Every response shape is validated against the request's `req_id` and
  expected `chunk_offset`; a reordered or duplicate reply marks the peer
  lost before any install. Retries request only the unfinished suffix.
- Worker connects are bounded and cancellation-aware; `stop` covers lock
  acquisition and the join under one deadline and detaches with a warn;
  blocking fetches run outside the install lock with a cancellation
  recheck before any copy.
- `fault_around_chunks_installed` rides the handler stats, the control
  report and the host's fault-path totals. Page-channel `PROTO_VERSION`
  5 -> 6; coordinator `WIRE_VERSION` 34 -> 35 so hosts roll in lockstep.

Tests: streamed installation, suffix-only retry, corrupt and reordered
replies, one-chunk equivalence with `NeedAt`, window selection over random
seals (proptest), concurrent drain dedup, faulting-page-only wake,
overlapping and unrelated faults and the drain bypassing a stalled suffix,
cancellation during a pending connect, a worker that misses the stop
deadline, a stalled AltSource fetch under stop; the two-host KVM test
checks the counter is reported.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UHxL6gj4o8EvxYpgtwEWaM
@nikhilunni
nikhilunni force-pushed the feat/postcopy-fault-around branch from 081008e to 2174654 Compare October 8, 2026 00:03

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant