Repository navigation
Stream Planner tool calls from their own start and end, including Atomic's own tools - #314
Open
lavaman131 wants to merge 2 commits into
Open
lavaman131 wants to merge 2 commits into
lavaman131 wants to merge 2 commits into
Conversation
lavaman131
force-pushed
the
planner-live-streaming
branch
2 times, most recently
from
October 6, 2026 22:14
594ec93 to
17d4f44
Compare
lavaman131
added a commit
that referenced
this pull request
Oct 6, 2026
Fixes that only exist because several open PRs now share one tree: - Drop #290's duplicate ImageIcon and its LineIcon titles, which #287 removed. - Remove the import #302 left unused after #276's inline code. - Renew design-contract pins for files more than one PR edited, and record #290's new icons and #295's face colours as reviewed exceptions. - Find icons by path rather than the titles #287 removed, and faces by the label #295 gives them. - A tool that returns after its turn was stopped ends as failed, so #314's early completion keeps #272's stopped-question outcome. Assistant-model: Claude Opus 5.5 Assistant-workflow: inline Assistant-verification: typecheck passed: bun run types Assistant-verification: dprint/oxlint/design checks passed: bun run ci Assistant-verification: bun test passed: 3860 pass, 0 fail Co-authored-by: Alex Lavaee <lavaman131@github.com>
…mic's own tools (#306) A Planner turn looked frozen: a tool row appeared only once the model's input was complete, and settled only when the harness finished the whole model step. Under HARNESS=atomic, Atomic's builtin, extension, workflow and subagent tools never reached the chat at all. If they had, the fixed PLANNER_TOOL_NAMES check would have aborted the turn. Server and protocol only. The chat presentation is Maggie's design work (#296), so apps/web is unchanged, and every wire addition is optional: - Activity.startedAt (server time when the tool's input began), Activity.refused, and Turn.startedAt in milliseconds. A client that ignores them keeps working. - A row opens on tool-input-start and settles when the tool itself finishes. Host tools report through a wrapper on execute, because HarnessAgent buffers results until step end. Partial output updates the running row. - Args and results are truncated and secrets masked. read_reference output stays private. - The Atomic adapter turns tool_execution_* and tool-call input streaming into stream parts for tools Atomic runs itself. A full session's allowed set is its running turn's active tools. copilot-sdk and pi still fail closed on any tool outside their fixed set. - Running tools are marked failed when a turn ends or a room is restored, and a mid-turn reconnect restores the turn and running rows from history. - Reasoning is deliberately not sent to clients. It can quote read_reference content, and showing it is a design decision for #296. Refs #306 Assistant-model: Claude Opus 5.5 Assistant-workflow: goal (run 1e1a9fe2-7263-4372-9eb3-31c679ccb69e), finished inline Assistant-duration: 3h15m converged, estimated 2h30m Assistant-verification: bun test passed: 3685 tests, 3 skip, 0 fail (translate tool-input/settle/boundary/reasoning-privacy cases, Atomic tool_execution_* and input-streaming cases) Assistant-verification: bun run types passed: all workspaces Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings) Assistant-verification: playwright e2e passed: bun run e2e e2e/harness.e2e.ts e2e/smoke.e2e.ts e2e/responsive-activity.e2e.ts, 28 passed (wire tests: tool row opens before text and settles at its own ~2.5s end inside a 7.5s turn; reconnect restores the running row and turn start) Assistant-verification: real GitHub QA passed: githubnext/pdd-demo with HARNESS=atomic on the final build; Atomic's builtin Bash showed live and settled at 4.9s for `sleep 4`; no boundary failure in the server log User-preference: Don't overlap teammates' in-flight UI work; leave chat, composer and design changes to Maggie's design track (#296) User-preference: Don't show Planner reasoning where it conflicts with Maggie's design Co-authored-by: Alex Lavaee <lavaman131@github.com>
…306) dev/integration carries #270, which withholds bash from the Planner's own turns, so its stub reads a file instead. Pick the tool from the stream's dynamic parts rather than naming bash, so the same tests hold on main and on top of #270. Assistant-model: Claude Opus 5.5 Assistant-workflow: inline Assistant-verification: bun test passed: apps/server/src/harness/atomic/full.test.ts (18 pass) on main's base Co-authored-by: Alex Lavaee <lavaman131@github.com>
MaggieAppleton
force-pushed
the
planner-live-streaming
branch
from
October 7, 2026 10:57
17d4f44 to
3afbf32
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A Planner turn looks frozen in Chat. A tool row only appears once the model has finished writing its input, and only settles when the harness finishes the whole model step, so its duration covers the step instead of the tool. Under
HARNESS=atomic, Atomic's builtin, extension, workflow and subagent tools (bash, read, workflow, subagent…) never reach Chat at all. If they did, the fixedPLANNER_TOOL_NAMEScheck would abort the turn.This PR fixes the data side only. The chat presentation belongs to @MaggieAppleton's design work (#296, #162), so
apps/webis unchanged and every protocol addition is optional.What changes
packages/protocol/chat.d.ts, all optional):Activity.startedAt(server time when the tool's input began),Activity.refused, andTurn.startedAtin milliseconds. A client that ignores them keeps working.apps/server/src/chat/service.ts):tool-input-start.HarnessAgentbuffers host-tool results until step end, soagents.tswrapsexecuteto report completion straight away, and the buffered part then finds the row already settled.chat/activity.ts): args and results are truncated to 4,000 characters, with token, password and API-key fields andghp_/github_pat_/sk-/Bearervalues masked.read_referenceoutput stays private.harness/atomic/adapter.ts,session.ts,full.ts):tool_execution_start/update/endand tool-call input streaming become stream parts for tools Atomic runs itself. Nested subagent calls stay under their parent row.copilot-sdkandpistill fail closed on any tool outside their fixed set (regression test).read_referencecontent that is private to the Planner, and whether to show it is a design call for Show Chopin's work as live stages with inspectable tools #296.Overlap with #296
Both PRs touch
Chat.Turninchat.d.ts, where #296 addsentryOffset, and the twochat.turn = {…}literals inservice.ts, which this PR replaces withbegin(handle). Both also append tests toe2e/harness.e2e.ts. The changes don't conflict in meaning. I suggest merging #296 first; I'll then rebase and addentryOffsettobegin().Verification
Real GitHub,
githubnext/pdd-demo,HARNESS=atomic, final build, withmain's UI unchanged. Atomic's builtinBashrow appears live, then settles at 4.9s forsleep 4: its own duration, not the turn's. The server log has no boundary failure.bun run e2e e2e/harness.e2e.ts e2e/smoke.e2e.ts e2e/responsive-activity.e2e.ts: 28 passed. The new wire tests record WebSocket frames: the busy state comes before any text, the tool row settles at its own ~2.5s end inside a 7.5s turn, and a reload mid-turn restores the running row and turn start from history without reasoning.bun test: 3685 pass, 3 skip, 0 fail. This covers translate cases for input streaming, own-end settlement, boundary failures and reasoning privacy, plus the Atomictool_execution_*and input-streaming adapter cases.bun run typesandbun run cipass, with no new Impeccable findings.Refs #306
Assistant-workflow: goal (run 1e1a9fe2-7263-4372-9eb3-31c679ccb69e), finished inline
Assistant-duration: 3h15m converged, estimated 2h30m
Assistant-verification: bun test passed: 3685 tests, 0 fail
Assistant-verification: playwright e2e passed: harness, smoke, responsive-activity, 28 passed
Assistant-verification: bun run types and bun run ci passed
Assistant-verification: real GitHub QA passed: pdd-demo, HARNESS=atomic, builtin Bash live and settled at 4.9s for sleep 4, no boundary failure
User-preference: Don't overlap teammates' in-flight UI work; leave chat, composer and design changes to Maggie's design track (#296)
User-preference: Don't show Planner reasoning where it conflicts with Maggie's design
Co-authored-by: Alex Lavaee lavaman131@github.com