Repository navigation
feat(adr-0054): harness special messages — file-change diffs + interactive AskUserQuestion - #381
Conversation
Proposed. Specs two additive harness surfaces, grounded in a protocol spike pinned against claude 2.1.181 (ten empirical findings recorded in the Decision): - Flavor A — render-rich FileChanged events so Write/Edit/MultiEdit render as diffs in the UI (Write = all-green new content; Edit/MultiEdit = red/green before/after). Additive wire + render, no invocation change. - Flavor B — interactive AskUserQuestion via a PreToolUse hook + an out-of-band unix socket (harness = server, hook = client). Always-defer/resume: a question ends the turn (VM idle-evictable), the answer arrives via AnswerQuestion and resumes the run through the existing ensure_active path (claude -p "" --resume re-fires the deferred tool, tool_use_id-stable, no transcript pollution). Wire additions are append-only over the bincode positional wire (Answers is a BTreeMap<String, Vec<String>> — no untagged enums, which panic under bincode 1.x). Status: Proposed — stays so until code lands. Two build phases: FileChanged (no invocation change) then interactive (drops --dangerously-skip-permissions for the auto-allow hook; the VM remains the isolation boundary). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…K audit) Refine Flavor B (interactive AskUserQuestion) with the empirical evidence that closes the resume seam findings #9/#10 only touched through discrete `-p` invocations, not the harness's persistent streaming process. - Add findings #11-13 (claude 2.1.183 + first-party @anthropic-ai/claude-agent-sdk 0.3.183): feeding the live process re-infers a NEW tool_use_id (can't answer a deferred tool); streaming-mode `--resume` re-fires the deferred hook on startup, id-stable, no stdin (so no `-p ""` trick); the Agent SDK confirms there is no in-process defer-continuation (canUseTool's PermissionResult is allow|deny only — cannot defer; the control protocol exposes only interrupt/stop_task). So kill -> respawn-with-`--resume` is the canonical path, matching the SDK. - Replace the placeholder `start_resume_run()` with the concrete path: stash into the run_engine-owned answers_in_hand, sigint_child, return a new SessionOutcome::ResumeForAnswer (re-using the existing respawn loop, not a crash), plus a continuation TurnState on the resumed startup so the re-fired output isn't dropped by the no-in-flight-turn branches. - Correct answers_in_hand scope (it spans the respawn) and the answer-delivery idempotency consequence (duplicate delivery is structurally safe, but at-least-once delivery is NOT yet provided — flagged for Phase 2). - Strengthen the canUseTool rejection: it cannot express defer at all, which is decisive, not merely lower-risk. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
Add the harness-agnostic interactive-question surface (Flavor B) ahead of
the harness/host/coordinator wiring that constructs it.
- `Question { question, header, multi_select, options }` + `QuestionOption
{ label, description }` + `type Answers = BTreeMap<String, Vec<String>>`.
`multi_select` is `#[serde(rename = "multiSelect")]` so the hook-bridge
can deserialize Claude's `tool_input.questions` payload straight into it.
Answers are a uniform `Vec<String>` (1-element = single-select) — NOT an
untagged `One|Many` enum, which would panic on the positional bincode wire.
- Trailing `HarnessEvent::UserQuestion` (11) / `QuestionAnswered` (12) and
`HarnessCommand::AnswerQuestion` (7) — append-only, existing indices
unchanged. Wired into `kind()` and `tool_call_id()`.
- wire_golden: golden + variant-index rows for all three, including a
mixed-arity `Answers` sample (a multi-element vec AND a 1-element vec in
one map) so any regression to a `deserialize_any` encoding fails loudly.
Regenerated corpus adds only the 3 new .bin files; no existing golden
changed (verified wire-safe).
Indices follow landing order (Phase 2 before Phase 1): the ADR's
documented 11/12/13 assumed Phase-1-first; FileChanged becomes 13 when
Phase 1 lands. ADR text updated in the closeout commit.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…hook + resume Make a deferred AskUserQuestion answerable. A PreToolUse hook defers the tool (ending the turn so the VM idle-evicts), the harness emits a durable UserQuestion card, and the user's answer drives a `claude --resume` re-spawn that re-fires the deferred tool id-stable and answers it (ADR 0054 findings #9/#11/#12 — the live process must NOT be fed; it would re-infer a new tool_use_id and never answer the deferred call). - `hook-bridge` subcommand (git/busybox argv dispatch in `main()`, before clap): the transient PreToolUse client. Ordinary tool → `allow` (the `--dangerously-skip-permissions` replacement); AskUserQuestion → socket round-trip, then denormalize the canonical Answers to claude's `updatedInput.answers` (bare string single-select / array multiSelect, finding #6) — the claude arity quirk stays isolated here. - `hook_server`: a `UnixListener` bound once before the first spawn (bind-before-spawn), shared across respawns; NDJSON one line each way; verdict = answer-in-hand (consume) → answer + QuestionAnswered, else defer + UserQuestion. Client/server share the serde types (no drift). - `SessionOutcome::ResumeForAnswer`: the `AnswerQuestion` arm stashes the answer into the run_engine-owned `answers_in_hand` (Arc<tokio Mutex>, survives respawn; touched by both the cmd arm and the accept task), SIGINTs claude, and `break`s (NOT `return` — the reap block must run). ResumeForAnswer respawns with `--resume` like Respawn but resets the fast-crash budget and never emits the abnormal-exit System message. - Continuation turn: on a respawn where `answers_in_hand` is non-empty, startup establishes a TurnState (fresh run_id, RunStarted with no prompt_id, no user-echo) so the re-fired tool_result + the model's continuation are captured, not dropped by the no-in-flight-turn branches. - Invocation: drop `--dangerously-skip-permissions`, add `--settings` (written at startup from `current_exe()` so the hook command is this binary's own path — correct on FC / VZ / Process), inject `ENGRAM_HOOK_SOCK` (finding #7). `IS_SANDBOX=1` kept as the root-check escape hatch (independent of the dropped flag; flagged in the ADR for an empirical re-check that claude still starts as root under the hook). Tests (run in CI — harness tests are Linux-gated, can't run on macOS dev): socket round-trip (defer→answer, with UserQuestion/QuestionAnswered), multi-question one-roundtrip, idempotent duplicate-refire, denormalize single/multi, and an engine-level answer→ResumeForAnswer→continuation-turn cycle via a fake claude. The socket failing to bind is non-fatal (ordinary tools still allow; only AUQ degrades to always-defer), so engine tests are unaffected by the test env's `/workspace` writability. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…nce replay
Route the user's answer to the in-guest harness and guarantee delivery
across a connection bounce, mirroring the prompt replay buffer.
- `HarnessHub::answer_question(sandbox, tool_call_id, answers)` sends
`HarnessCommand::AnswerQuestion`. An answer resumes an idle/evicted
session — precisely when the post-resume attach race bites — so it shares
`send_prompt`'s attach-wait via a new `acquire_cmd_tx_waiting` helper
(extracted from `send_prompt`, same lock-order discipline).
- `undelivered_answers` replay buffer (leaf lock, twin of
`undelivered_prompts`): record on send, replay on every (re)attach in
`drive_attached`, retire on `QuestionAnswered{tool_call_id}` in
`reader_loop`, drop on `unbind_session`. At-least-once: a dropped answer
is auto-redelivered; the harness's resume is idempotent (the deferred tool
yields exactly one tool_result), so re-delivery is a no-op.
- harness-noop: handle the new trailing `AnswerQuestion` command arm
(exhaustiveness; noop never asks, so it's a no-op).
Test: answer_question delivers the frame, buffers it un-confirmed, and a
QuestionAnswered event from the harness retires the buffer entry.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…event passthrough Complete the answer path from the app-gRPC edge down to the host, and pass the two new harness question events through the session event log. - SessionEvent: `HarnessUserQuestion`/`HarnessQuestionAnswered` variants + `from_harness` arms + `kind()` (`user_question`/`question_answered`). Events are opaque JSONB + a kind string, so NO migration is needed. - `answer_question_core` (api/prompt.rs): mirrors `send_prompt_core` — reuses a new `ensure_active_and_resolve` helper (auto-resume + the mid-move HOLD, extracted from send_prompt_core so both share it) and forwards `HostClient::answer_question`. Emits NO user-echo: the harness's own `QuestionAnswered` is the durable answered record. - App-gRPC: `rpc AnswerQuestion` + request/response + a `StringList` map-value wrapper (proto3 maps can't hold `repeated`); handler unwraps the map into the canonical `Answers` and calls the core. - HostClient trait: defaulted `answer_question` (no-op for harness-less test fakes; HostRegistry / LocalHostClient / GrpcHostClient override). Signature spells out `BTreeMap<String, Vec<String>>` rather than `engram_harness_proto::Answers` to keep engram-core free of a dep cycle. - Remote path (production): host-control proto gains `rpc AnswerHarnessQuestion` + request (+ StringList); GrpcHostClient builds it; the host-agent gRPC server unwraps it and calls `LocalHostClient::answer_question` → the hub. So an answer reaches the harness whether the owning host is local (--mode=all) or remote (WS/gRPC). Test: from_harness maps both events under their stable kinds. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…nerated stubs Expose the coordinator's new `SessionService.AnswerQuestion` RPC through the orchestrator so the backend answer loop is reachable over HTTP (the web component is a follow-on branch). - Regenerate the app-contract TS stubs (`buf generate`) for the Group D proto change — `AnswerQuestionRequest`/`AnswerQuestionResponse`/`StringList` in both orchestrator and web `src/gen` (keeps the proto↔gen drift check green; these are additive type stubs, not the deferred web UI). - policy-map: `SessionService.AnswerQuestion` → owner-scoped `prompt` capability (answering a question is the same "drive a session you own" permission as sending a prompt). The descriptor-driven passthrough + SURFACE auto-include it from `SessionService.methods`; no handler code. typecheck + the full orchestrator test suite (incl. the auto-driven passthrough conformance + authz matrix) pass. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
Flip Status to Accepted with the backend-first commit chain, and reconcile the design doc with what landed: - Wire-surface indices follow LANDING order (Phase 2 before Phase 1): UserQuestion = event 11, QuestionAnswered = event 12, AnswerQuestion = command 7; FileChanged becomes event 13 when Phase 1 lands (clean trailing append, no reserved placeholder). - The `--settings` file is GENERATED AT STARTUP from `current_exe()`, not baked — the hook command needs the running binary's own absolute path, which varies across FC / VZ / Process. The binary is the one baked artifact; deriving settings from it is the single source of truth. - At-least-once answer delivery is now PROVIDED (the preferred option): the host-agent gained an `undelivered_answers` replay buffer retired by QuestionAnswered. Consequence bullet updated from "not yet provided". - Phase 2 build-out bullet updated to the actual landed surface (host-agent replay, coordinator ingress + event passthrough, orchestrator passthrough); web question component explicitly deferred to a follow-on branch. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
`answer_question_resumes_into_continuation_turn` raced: the fake claude's FIRST invocation had to `touch` a marker file before being SIGINT'd by the `AnswerQuestion`, and the resumed invocation only emitted output if that marker existed. Under load the SIGINT won the race, the marker was never written, the resumed invocation emitted nothing, and the test hung forever (surfaced running the Linux-gated suite in a container 40× — it passed at 0.6s on a quiet box but hung on a loaded one). Fix: make the fake stateless — it emits the same one-turn output (assistant text + `result`) on EVERY invocation, then blocks on stdin. The HARNESS state, not the fake, now decides the output's fate, with no test-side race: the first (idle) spawn has no in-flight turn so the output is dropped (only the startup `Idle` surfaces); after the `--resume` respawn the continuation `TurnState` captures it. Verified 40/40 iterations + full suite green on aarch64-linux. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
…ocker The `engram-harness-claude` hook/socket/resume suite is `#[cfg(target_os = "linux")]`, so it can't compile or run on a macOS dev box — only in CI. `just test-linux [pkg]` closes that gap: on Linux it's a plain `cargo test`; on macOS it builds + runs the crate's tests inside a `rust:bookworm` container (native arch on Apple Silicon — no emulation), erroring clearly if Docker is absent. The host cargo registry is mounted so crates aren't re-downloaded, and the target dir is a named volume (Docker-VM native fs — fast incremental rebuilds, persists across runs) kept separate from the macOS `target/`. Defaults to `engram-harness-claude`; pass any Linux-gated crate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0113QBWNWDZxGjcttD7p8Q7z
Render the deferred `AskUserQuestion` round-trip in the transcript. The
harness emits a durable `user_question` card (and a `question_answered` when
the user's reply re-fires the tool on `--resume`); the web now turns that
into an interactive form and POSTs the answer back via
`SessionService.AnswerQuestion`, completing the loop the backend branch
opened.
- events/sse: add `user_question` + `question_answered` to the SessionEvent
union and the per-kind SSE listeners (the orchestrator relay is
kind-agnostic, so no relay change).
- buildMessages: render `user_question` as a `UserQuestionCard` in the system
("harness register") lane, and DEDUP the generic `AskUserQuestion` tool part
by tool_call_id — the harness emits both. The answer arrives in the LATER
resume run, so it's folded onto the card via a pre-scan. A question-only run
no longer synthesizes a stray empty assistant bubble (the run receipt
attaches to the run's real assistant, or is skipped).
- UserQuestionCard: a form (single-select = radios, multi-select = toggles,
each option showing its label + description) whose submit POSTs answers
keyed by QUESTION TEXT (the wire contract — finding #8), values = selected
option labels wrapped in StringList. Optimistic receipt until the
authoritative `question_answered` confirms; rolls back on send failure.
- SessionThread owns the `answerQuestion` mutation + optimistic state, exposed
via `QuestionActionsContext` (mirrors composer-actions).
Tests: buildMessages cases (card, dedup, no-empty-bubble, cross-run answer
fold) + a SessionThread integration test (select -> submit -> request shape ->
optimistic receipt) through a router transport. tsc/oxlint/oxfmt/vitest + vite
build all green (107 web tests).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Append the rich file-change event for Phase 1 (Flavor A). `FileChanged {
run_id, tool_call_id, path, change }` is the next trailing HarnessEvent
(bincode variant 13, after QuestionAnswered=12) + `kind() = "file_changed"`
+ the `tool_call_id()` arm. `change` is a `FileChange` enum (`Write {
content }` / `Edit { hunks: Vec<EditHunk> }`) with `EditHunk { old, new }`,
and a 64 KB `MAX_FILE_CHANGE_BYTES` per-string budget.
Two deviations from the ADR's original sketch, both forced by the wire and
caught by wire_golden:
- `FileChange` is EXTERNALLY tagged (`#[serde(rename_all="snake_case")]` →
`{"write":…}` / `{"edit":…}`), NOT `#[serde(tag="op")]`. It rides bincode
inside `FileChanged`, and an internally-tagged enum uses `deserialize_any`,
which bincode rejects (`DeserializeAnyNotSupported`) — the exact rule this
ADR cites for `Answers`. The golden decode failed loudly on the `tag="op"`
first cut. The JSON the coordinator re-emits over SSE stays clean.
- `path` is hoisted onto the EVENT (both ops have one), so `FileChange`
carries only op-specific payload instead of repeating `path` per arm.
wire_golden pins both new variants (edit w/ a replacement + a pure-insertion
hunk, and a write) + variant index 13; regen added only the two new
golden .bin files (no existing bytes changed → wire-safe append).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
… edit Recognize `Write`/`Edit`/`MultiEdit` in the stream-json observer and emit a `FileChanged` correlated by `tool_use_id`. `parse_file_change` maps the tool INPUT to a normalized `(path, FileChange)` (Write→content, Edit→one hunk, MultiEdit→many; each inner string clipped to MAX_FILE_CHANGE_BYTES). The parse is stashed in a per-turn `pending_file_changes` map (on TurnState, alongside `current_message_id`) when the `tool_use` is seen, and drained on the matching `tool_result`: - success (`is_error:false`) → emit FileChanged (after ToolCallCompleted), - failure → drop the stash, emit NO FileChanged (truthful — never a phantom diff for a failed edit). The generic ToolCallStarted/Completed still fire; the web dedups by tool_call_id and renders the rich diff in place of the generic card. Tests (run via `just test-linux` on macOS): a successful Edit emits FileChanged with the hunk + retires the pending entry; a failed edit emits none and still drops the stash; a non-file tool stashes nothing. Full Linux-gated suite green (20 passed); clippy clean on the linux target. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Map `HarnessEvent::FileChanged` to a `SessionEvent::HarnessFileChanged` passthrough variant (`from_harness` arm + `kind() = "file_changed"`), carrying `path` + `change` opaquely. Like every other passthrough event it's JSONB on the existing `session_events` column — no migration. Flows through `harness_event_sink` → `session_events` → SSE unchanged. Test asserts the stable `file_changed` kind and that `FileChange` serializes externally-tagged (`change.edit.hunks` is an array) — the snake_case shape the web discriminates on. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Render the `file_changed` event as an inline diff card, completing Flavor A
end to end.
- events/sse/types: add the `file_changed` event + the `FileChange` shape
(externally-tagged `{write}` / `{edit}`) and the per-kind SSE listener.
- buildMessages: pre-scan `file_changed` by tool_call_id; a Write/Edit/
MultiEdit tool call that produced one is re-rendered as a `FILE_CHANGE_TOOL`
part carrying `{path, change}` (still tallied as an edit) instead of the
generic card. A FAILED edit emits no file_changed, so it keeps its generic
error card. Reuses the AskUserQuestion dedup pattern.
- FileChangePart: a collapsed row — path + `+N −M` counts computed eagerly
with jsdiff (cheap, no Shiki) — expandable to a Pierre (`@pierre/diffs`)
`FileDiff` built from before/after via `parseDiffFromFile` (Write = empty →
content; Edit = joined old → joined new). The Pierre+Shiki renderer is
`React.lazy`'d into its own chunk so the main bundle barely grows and
Shiki's WASM/grammars load only on expand (verified: PierreDiff is a
separate 420 KB async chunk; index bundle unchanged within ~11 KB).
- Registered in thread.tsx beside ShellToolPart.
deps: `@pierre/diffs` (Apache-2.0) + `diff` (jsdiff, for eager counts).
Tests: buildMessages swap/dedup/failed-edit-keeps-generic + FileChangePart
collapsed row (path + counts, diff not mounted). tsc/oxlint/oxfmt/vitest
(112) + vite build all green.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Record that rich file-change rendering shipped: `e32e8311` (proto FileChanged wire types) · `a6870338` (harness emit on successful Write/Edit/MultiEdit) · `b2b54b81` (coordinator file_changed passthrough) · `0a5e3e4b` (web Pierre FileChangePart). Update the Status block + the Flavor A schema sketch to the as-built shape, and document the two wire-forced reconciliations: FileChange is externally-tagged (not `tag="op"`, which panics on bincode decode), and `path` is hoisted onto the event. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
…ate-past claude-sonnet-4-6 does not reliably end a deferred AskUserQuestion turn as terminal_reason=tool_deferred. On the 2nd question in a session (~1/8 reproduced; opus-4-8 0/4) it records hook_deferred_tool, then runs an extra inference concluding "internal error retrieving your answer" and ends the turn `completed` — never suspending, defeating the defer->resume wait. subtype/is_error are identical to a clean defer, so the discriminator is terminal_reason + deferred_tool_use (the Agent SDK contract). Fix: detect_result_marker now parses terminal_reason; translate_jsonl tracks pending (no-tool_result) AskUserQuestion tool_use_ids per turn and suppresses any assistant text/chunks emitted while one is pending, so the hallucinated reply never reaches the UI. The UserQuestion card the hook emitted stays the source of truth (session remains awaiting-answer). The session loop logs the narrate-past for observability. TDD: 3 new failing tests + 2 regression guards (answered AUQ and same-message preamble are NOT suppressed). Found via session bc04ed42 + a two-question repro probe. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Propagates the AskUserQuestion narrate-past fix (a56375b) up the stack. Resolved translate_jsonl conflicts so TurnState, the fn signature, and every call-site carry BOTH pending_file_changes (Flavor A) and auq_pending (the fix); the cfg(linux) test call-sites updated to 7 args. Verified on Linux: 25 harness tests pass, clippy -D warnings clean, fmt clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
…d questions (Part B) Part A suppressed the hallucinated reply and kept the UserQuestion card, but the card was a dead end: VERIFIED that after a narrate-past, a --resume does NOT re-present the abandoned deferred tool (hook fires 0x; claude misremembers the prior answer), so an answer stashed for it never lands and the card hangs forever. Part B (try-then-fallback): when an answer-resume continuation turn ENDS with the answer still un-consumed in answers_in_hand, the deferred tool was abandoned -> deliver the answer as a fresh user message (claude asked conversationally, so a plain message is what it awaits) and emit QuestionAnswered to mark the card. No pre-classification needed and it survives idle-eviction between question and answer. TurnState.is_answer_resume flags continuation turns; fallback_answer_message renders the answer map. TDD: rewrote answer_question_resumes_into_continuation_turn -> answer_resume_unconsumed_falls_back_to_user_message (responsive fake claude); 23 tests pass, clippy/fmt clean on Linux. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
…ript before resume
After a PreToolUse `defer`, claude nondeterministically runs a continuation
inference that hands the model `tool_result{is_error:true, "[Tool result missing
due to internal error]"}` for the deferred tool (upstream bug
anthropics/claude-code#64389). Part A already suppresses that hallucinated text
from the UI, but it still persists in claude's own `.jsonl` transcript — so on
`--resume` the model reads its own stale "internal error / I'll ask again" text
and re-asks regardless of the real answer we deliver.
Part C records the message-ids Part A suppresses (a respawn-surviving set) and,
in the `ResumeForAnswer` gap (claude is dead, transcript quiescent), deletes
exactly those assistant lines from the transcript (`scrub_transcript`: match by
message.id, re-link parentUuid, atomic temp+rename) so the resumed transcript is
byte-equivalent to a clean defer. Fail-safe: any miss/parse error skips the
scrub and Part B's answer-as-message fallback still delivers the answer.
Wire-level alternatives (rewrite the is_error result, drop the continuation via
a synthetic empty response, inject 5xx/400) were all tried and rejected — each
leaves a residue that re-breaks resume; only removing the persisted text reaches
the clean-defer state.
Validated: 10/10 local prototype, 28/28 Linux tests (3 new scrub tests via
`just test-linux`), and end-to-end in-stack on a real narrate-past (scrub fired
`removed=1`, the answer landed, no re-ask).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
Brings the ADR 0054 stack (Phase 1 file-diffs + Phase 2 AskUserQuestion + Part C transcript scrub) up to date with main (ADR 0052 Phase 1b/1c/3/4 — queued/streamed messages, control_request interrupts, warm teleport; ADR 0056 integration assets). Conflict resolutions (union both sides; nothing dropped except main's intentional ADR 0056 retirement of PullRequestOpened): - engram-harness-claude/src/main.rs: turn-loop unions ADR 0054's defer/ answer-resume + Part C scrub with main's Phase 3 interrupt machinery (interrupt_deadline, control_request); test module is the union of all 31 tests (HEAD's AUQ/FileChanged/scrub + main's interrupt/connection-bounce). - engram-coordinator/src/state.rs: kept ADR 0054 events (HarnessUserQuestion/QuestionAnswered/FileChanged); replaced the retired PullRequestOpened with main's IntegrationAsset (ADR 0056). - web buildMessages.ts/.test.ts: unioned AUQ + file-change rendering with main's stable-assistant-ids + queued-prompt logic; PR rendering now flows through integration_asset (forge/pull_request). Generated *_pb.ts and pnpm-lock.yaml regenerated. Verified: harness 31 tests pass on Linux (just test-linux), coordinator compiles, web tsc + vitest (120) pass, orchestrator typecheck clean. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01RnW5u5MECthL62FeyB2hrs
…le-fire at the hook With stdin held open, claude occasionally runs a continuation after deferring an AskUserQuestion and re-calls the tool with a fresh message/tool_use id, yielding two deferred AUQs and two question cards for one logical question (observed in sessions b63a8cbf, 76207d3c). This is the tool-call sibling of the narrate-past text (anthropics/claude-code#64389); the turn still ends tool_deferred, so the terminal_reason narrate-past guard never catches it. Capture it at the hook server. serve() now owns a per-run set of run_ids that have already deferred an AUQ; the first AUQ in a run defers + shows a card as before, and any subsequent AUQ in the same run is DENIED (permissionDecision "deny" + reason). Denying resolves the duplicate as a tool_result instead of letting it become a second deferred tool, so on --resume there is exactly one deferred AUQ to re-fire — no orphan card, no tool_use_id re-mint (the reason scrubbing a mid-chain duplicate fails). HookVerdict gains Deny{reason}; the bridge gains a print_deny arm. Validated: claude settles on the deny without looping and the turn still ends tool_deferred (local held-open-stdin repro, 11/11); 32 harness tests pass including a new hook_denies_duplicate_auq_in_same_run; no regression across 30 in-stack turns (the bug did not reproduce on the freshly-baked image, so a live collapse was not captured). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HAGpc8yyuYkEUpYf53zVwz
… the #64389 duplicate AUQ The #64389 stdin double-fire also re-CALLS AskUserQuestion for one logical question (the tool-call sibling of the narrate-past text), producing a second question card. It shows up two ways: WITHIN a run, and — distinctly — ACROSS a run boundary (the deferred turn ends tool_deferred, then claude, still resident on held-open stdin, spontaneously starts a new run and re-asks; session 709008e4). The prior cut (218b0e5) denied the duplicate, but a deny is itself transcript poison: it writes a "stop and wait" tool_result the model RETRIES after the real answer lands (re-ask + second card, session 691828f6), and its per-run keying misses the cross-run shape entirely. Replace it with a SESSION-level invariant — at most one outstanding (carded, unanswered) question per session: - the hook owns `question_outstanding` (set when it cards the first AUQ, cleared when that answer is delivered); while set, every further AUQ — same run or a later one — DEFERS with no card (no deny poison) and its tool_use_id is recorded in `duplicate_auq_ids`. - `scrub_transcript` is extended to also drop any assistant message carrying a duplicate `tool_use` (matched by that id, so it reaches a cross-run sibling), leaving the resumed transcript with exactly one pending question. - the duplicate's `ToolCallStarted` is suppressed at the emit site (a transient phantom tool-call otherwise flickered in the UI). Removes the `HookVerdict::Deny` machinery (DUPLICATE_AUQ_REASON, print_deny) — a clean break, no compat shim. Validated: 33 harness tests pass (incl. hook_dedups_duplicate_auq_session_wide, scrub_removes_duplicate_auq_tool_use); clippy -D warnings clean on the aarch64-unknown-linux-musl target; live repro (session 34eccff9) hit BOTH the within-run and cross-run double-fire in one run — one card, scrub removed=1, clean answer + haiku, zero re-ask. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HAGpc8yyuYkEUpYf53zVwz
Resolve the single conflict in SessionProfileEditor.tsx: main replaced the capsText <Textarea> with the catalog-driven <CapabilityPicker> (ADR 0057 D1); this branch's incidental reformat of the now-deleted textarea is dropped in favor of main's version. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
|
The latest Buf updates on your PR. Results from workflow CI / buf (pull_request).
|
| if !stale.is_empty() { | ||
| if let Some(s) = stdin.as_mut() { | ||
| let text = fallback_answer_message(&stale); | ||
| let ft = start_turn( | ||
| evt_tx, s, cli, None, &text, current_run_id, | ||
| ) | ||
| .await; | ||
| for (tool_call_id, answers) in stale { | ||
| emit( | ||
| evt_tx, | ||
| HarnessEvent::QuestionAnswered { | ||
| run_id: ft.run_id.clone(), | ||
| tool_call_id, | ||
| answers, | ||
| }, | ||
| ) | ||
| .await; | ||
| } | ||
| turn = Some(ft); | ||
| } |
There was a problem hiding this comment.
Bug: question_outstanding is never cleared on this Part B fallback path → the session-level AUQ dedup leaks and silently swallows every later question.
question_outstanding (the session-level "a card is awaiting an answer" flag) is set back to false in exactly one place — the hook's answer-consumption path (main.rs:567). But this Part B fallback is taken precisely when claude narrate-past's and abandons the deferred tool, so on --resume the hook never fires — line 567 is unreachable on this path. The answer is delivered as a user message and QuestionAnswered is emitted here, but the flag stays true for the rest of the session.
Consequence: after a single narrate-past fallback, every subsequent genuine AskUserQuestion in the session hits the duplicate branch (main.rs:583) → deferred with no card → silently swallowed (and later scrubbed from the transcript). That re-introduces the exact "user never sees the question, Claude gives up" failure this ADR set out to fix. Narrow trigger (needs the ~1/8 narrate-past plus a later question in the same session), but silent when it hits.
One-line fix — clear the flag here (question_outstanding is already a param of run_claude_session), before start_turn so a fresh question asked in response is still carded:
if !stale.is_empty() {
// ADR 0054: the card is now answered via the fallback; the hook never
// fired to clear it on this path, so clear the session-level flag here
// or every later AUQ is wrongly deduped as a duplicate.
*question_outstanding.lock().await = false;
if let Some(s) = stdin.as_mut() {
...The invariant becomes "the flag is cleared on every answer-delivery path," not just the hook path.
Test gap: hook_dedups_duplicate_auq_session_wide covers the hook clearing the flag, but answer_resume_unconsumed_falls_back_to_user_message never sets it, so this path is untested. Worth extending that test to assert the flag is false after the fallback (and that a follow-up question re-cards).
… B2b) (#397) ADR 0057 B2b (0c66a4a, "coordinator sources network+secrets from the policy; strip manifest") removed the network allow-list from the image manifest and now sources egress from the session policy, defaulting to deny-all when a session has no policy. It updated engram-cli to send an allow-all policy on its debug creates but missed the e2e Driver, which still sent an empty integration_policy_json — so every e2e session now runs under deny-all egress. Only e2e_claude_with_bogus_key_surfaces_anthropic_auth_error visibly fails: it's the one test that needs egress (it expects Claude to reach api.anthropic.com and surface the 401 "Invalid API key"), so under deny-all Claude can't reach the API, never gets the 401, and the test times out at 180s. The other suite tests make no external calls and pass either way. This sat undetected because the e2e lane is path-gated and was skipped on every ADR 0057 merge; #381 was the first change since to re-trigger it. Fix: the Driver's three create methods send the same allow-all network policy engram-cli uses ({"network":{"default":"allow"}}), via a shared const. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…n provider icons (#401) Closes out the web redesign on the merged stack. - Launch page (`/launch`, spec §H) — the developer's full-page new-session picker: a profile-card grid (each card shows its provider tiles + power count), a task box, and the live "this session will be able to" receipt (per-provider read/write powers + reachable hosts, or "Run fully sandboxed"), reusing `derivePolicy`. Launches via the same CreateTask path as the quick dialog, then routes to the new session. Reachable from a "Launch" button on the Tasks page. - In-session provenance — the `integration_asset` timeline renderer now resolves the provider's identity via `useProviderIdentity` and renders its `ProviderTile` (instead of a generic puzzle icon / raw provider id), in both the PR card and the generic asset card. The pull-request renderer is now keyed on the asset kind (provider-agnostic), retiring the stale `forge/pull_request` switch. - `catalogToViews` — a member-safe `ConnectorView` builder from the member catalog alone (no admin `listConnectors`), for these member-facing surfaces. Tests: a Launch contract test (cards + receipt + CreateTask wiring); the #381 session-thread tests still pass (the asset renderer degrades to the monogram when the catalog is unavailable). Web lane: tsc -b clean, oxlint clean, oxfmt clean, vitest 144 pass (25 files), vite build OK. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Implements ADR 0054 end-to-end — a family of special harness messages: first-class, harness-agnostic events the web renders or interacts with richly, with no per-agent decoder (consumers read structured fields, not agent bytes). Two flavors.
Flavor A — rich file-change diffs (one-way)
Write/Edit/MultiEditused to render as opaque key/value blobs (the 1 KBargs_summarycap even destroyed the data a diff needs). The harness now recognizes those tools in claude'sstream-jsonand, on the successfultool_result, emits a typedFileChanged { path, change }event. The web swaps the generic tool card for a collapsed+N −Mrow (counts via jsdiff — cheap, no Shiki) that expands to a full Pierre (@pierre/diffs) diff — all-green forWrite, red/green hunks forEdit/MultiEdit— lazy-loaded into its own chunk so Shiki's WASM stays out of the main bundle. A failed edit emits noFileChanged, so it keeps its generic error card (never a phantom diff).Flavor B — interactive
AskUserQuestion(round-trip)Today
AskUserQuestionis dead: the CLI auto-denies it ~3 s after thetool_useand the user never sees the question. This wires it up:PreToolUsehook (the harness binary re-invoked ashook-bridge) replaces--dangerously-skip-permissions— ordinary tools auto-allow (explicit + auditable; the VM is still the safety boundary),AskUserQuestionbridges to the harness over a per-session unix socket (ENGRAM_HOOK_SOCK, bind-before-spawn).UserQuestionevent. The web renders aUserQuestionCard(radios for single-select, toggles for multi-select) and marks the session awaiting-input.AnswerQuestion→ensure_active(auto-resumes the FC snapshot if the session idled) → the harness stashes the answer and re-spawnsclaude --resume. The deferred tool re-fires id-stable, the hook returns the answer, and the run continues.QuestionAnsweredconfirms back onto the card.Why defer/resume, not a held permission
A question is human-paced and engrams is a density service, so pinning a hot VM to wait on a click is the wrong trade. The round trip is grounded in a protocol spike — 13 empirical findings pinned against claude 2.1.181 / 2.1.183 + the first-party Agent SDK 0.3.183 (recorded in the ADR). Decisively: the first-party SDK also continues a deferred tool only via a fresh
--resumespawn (canUseToolcan'tdefer; the live control protocol has no verb to continue a deferred tool), so this is the canonical path, not an engram limitation.Hardening for upstream claude bug #64389
A
PreToolUsedefernondeterministically (~1-in-6 locally; any model/tool) runs a continuation inference that poisons claude's transcript — a hallucinated "internal error retrieving your answer", or a duplicate question. Four layered mitigations:--resumeends with the answer still unconsumed (claude abandoned the deferred tool), fall back to delivering the answer as a fresh user message and mark the card answered — the question never hangs.scrub_transcriptdeletes the exact suppressed narrate-past lines from claude's.jsonl(match bymessage.id, relinkparentUuid, atomic temp+rename) so the resume starts byte-equivalent to a clean defer. Fail-safe: any miss falls back to Part B.Wire & delivery
HarnessEvents (UserQuestion=11,QuestionAnswered=12,FileChanged=13) andHarnessCommand::AnswerQuestion=7 are trailing, append-only over the positional bincode wire.wire_golden.rspins both the bytes and theu32variant index for each — including a mixed-arityAnswersmap so a regression to an untagged/deserialize_anyencoding fails loudly.Answers = BTreeMap<String, Vec<String>>(no untagged enums — they panic under bincode 1.x).#[serde(default)], no migration); an old coordinator degrades gracefully (drops the live NOTIFY, replays from PG).undelivered_answersalongsideundelivered_prompts, recorded onanswer_questionand retired onQuestionAnswered{tool_call_id}.AnswerQuestionis owner-scoped, byte-identical toSendPrompt(fail-closed).Layers touched
engram-harness-proto(wire types + golden) ·engram-harness-claude(the bulk — hook server +hook-bridgesubcommand + defer/resume engine + narrate-past suppression + transcript scrub) ·engram-host-agent(answer ingress + replay) · coordinator (event passthrough +AnswerQuestioningress reusingensure_active) · orchestrator (passthrough + authz) · web (FileChangePart/PierreDiff+UserQuestionCard+buildMessagesdedup bytool_call_id).Full design, the 13 findings, and alternatives considered:
docs/adr/0054-harness-special-messages.md.🤖 PR description generated with Claude Code