Skip to content

Keep the Planner while its workflows run, with Stop/Resume and a run card - #230

Merged
lavaman131 merged 21 commits into
mainfrom
atomic-planner-run-control
Oct 5, 2026
Merged

lavaman131 merged 21 commits into
mainfrom
atomic-planner-run-control

Conversation

@lavaman131

Copy link
Copy Markdown
Collaborator

Stacked on #215. Partly implements #221.

An atomic Planner session that starts an Atomic workflow now outlives its turn. Until now, Chopin destroyed the session when the turn ended, which left the run orphaned seconds after launch.

What changes

  • The session stays while it owns runs. An atomic Planner session is kept, with its owner binding and the loaded document, as long as it has live or paused workflow runs. Chopin tracks those runs through Atomic's workflow activity stream. The next turn reuses the session, so the Planner can still steer its runs over Intercom. The session is released when its runs end, when the owner binding ends, or when the document closes.
  • Stop and Resume. Stop Planner pauses the session's live runs through the session's own workflow tool, which returns a structured result instead of a dropped UI notice. Atomic can't pause a run whose only live stage is waiting on a question, so that run is quit instead. It stays resumable and its open question is withdrawn from Decisions. The new chat:resume message and the Resume Planner button resume paused runs.
  • Run card in Chat. Each run gets a card showing its name, status (running, waiting on Decisions, paused, finished, blocked, failed, or stopped), elapsed time, and each stage with its status and duration. A pulsing dot marks live work and stays still under reduced motion. Ended runs collapse to one line with the stages behind a toggle, and the transcript records how each run ended. A design-audit specimen covers the running, waiting, paused, blocked and finished states.
  • create_document outcomes. It now reports title-taken, and declares its idempotency-conflict, document-unavailable and validation-issue outcomes in its output schema. Before, strict MCP clients rejected those responses as schema violations.
  • Working directory. Without a checkout, the Planner's working directory moves from the shared tmp directory to a per-channel directory under Chopin's per-user state directory.

Verification

  • bun test: 1755 pass, 2 PostgreSQL skips, 0 fail.
  • bun run types and bun run ci pass, with no new Impeccable findings.
  • Manual agent-browser run against a local server, with the full atomic Planner running a planning workflow on a private demo repository:
    • the run card tracked the stages;
    • Stop paused the run and Resume continued it;
    • two Decisions questions were answered in the browser;
    • the workflow wrote its formal specification and plan, and the document reached revision 12.

Known gaps

Assistant-workflow: inline
Assistant-verification: bun test passed: 1755 pass, 2 skip, 0 fail
Assistant-verification: bun run types and bun run ci passed
Assistant-verification: agent-browser E2E passed: run card, Stop/Resume, Decisions answered, workflow wrote its spec and plan
Co-authored-by: Alex Lavaee lavaman131@github.com

@lavaman131
lavaman131 added this pull request to stack #220 September 30, 2026 17:13
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from d18de6e to 0931da6 Compare October 1, 2026 15:27
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch 2 times, most recently from 0931da6 to 1b76647 Compare October 1, 2026 15:42
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from 1b76647 to 24659d0 Compare October 1, 2026 16:18

@MaggieAppleton MaggieAppleton left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd fix three things before merging:

  • Save progress cards and finish messages before showing them, so a failed save cannot leave people seeing progress that disappears after a reload.
  • Keep the completion card and message when a quick job finishes before the Planner's reply ends.
  • Recheck cleanup after a reply finishes, and keep old cleanup from overwriting a newer helper's progress.

The small code suggestions passed focused regression tests. The save-ordering fix needs the card and message paths changed together; I've described the behavior and test needed in its comment.

Comment thread apps/server/src/chat/service.ts Outdated
Comment on lines +923 to +925
chat.runs = cards.length ? cards : undefined;
state(chat, server, room);
context.persist().catch(err => console.error("[chat] storing workflow runs failed:", err));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's save the new card and its finish message before showing either to people. Right now the message is sent by say() above, then the card is sent here, and only afterwards do we try to save them. If saving fails, people can see a job finish and then lose that update after a reload.

Please handle these updates in order, save the card and message together, then send both. Moving only state() would still send the message too early. Please add a test where saving fails and check that neither update was announced. This needs a coordinated change, so a one-line suggestion would leave part of the bug in place.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c29f101. A run report is now built as a candidate outside shared chat state and saved first: the save stores the candidate card and transcript together, and only when it succeeds are the card and finish message added to the chat and announced together. While a save is pending, or after it fails, nothing else that broadcasts chat state or greets a joining client can show them. Reports are handled one at a time, so a report that arrives during a pending save is computed from the state after it settles.

The new test fails the save and checks that neither update was announced; others cover overlapping reports, a state broadcast during a pending save, and a failed save of a run's final report, which is replayed on release so the run still ends as finished rather than "was stopped".

Assistant-verification: fail-before/pass-after passed: save-failure, overlapping-report, state-leak and failed-final-save tests; bun test 1928 pass, 0 fail on the rebased branch

Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Co-authored-by: Alex Lavaee lavaman131@github.com

Comment thread apps/server/src/chat/service.ts Outdated
Comment on lines +941 to +943
if (!runs || !(runs.active.length || runs.paused.length) || opened.binding.signal.aborted) {
return false;
}

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A quick job can finish before the Planner's reply is done. We return here without keeping its card, so people never see that it finished or failed. I reproduced the missing card. Let's record its last state even when there is no helper left to keep running. Please apply this with the companion finish-message suggestion above.

Suggested change
if (!runs || !(runs.active.length || runs.paused.length) || opened.binding.signal.aborted) {
return false;
}
if (!runs || opened.binding.signal.aborted) return false;
if (!runs.active.length && !runs.paused.length) {
publishRuns(context, runs);
return false;
}

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied in c29f101, together with the finish-message change: a quick job that ended before the Planner's reply finished keeps its last card and gets its finish message. The quick-job test fails without it.

Assistant-verification: fail-before/pass-after passed: quick-job test

Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Co-authored-by: Alex Lavaee lavaman131@github.com

Comment thread apps/server/src/chat/service.ts Outdated
for (let run of cards) {
let ended = ENDED_RUN[run.status];
let previous = before.get(run.id);
if (!ended || !previous || ENDED_RUN[previous]) continue;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the other half of the quick-job fix: a job we first hear about after it has ended should still get a finish message. The current check skips it because there is no earlier card. This suggestion allows that first message and still avoids repeating one for a job already recorded as finished.

Suggested change
if (!ended || !previous || ENDED_RUN[previous]) continue;
if (!ended || (previous && ENDED_RUN[previous])) continue;

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied in c29f101: a run first reported as already ended gets one finish message, and a run already recorded as ended never gets another. The first-seen-ended test fails without it.

Assistant-verification: fail-before/pass-after passed: first-seen-ended test

Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Co-authored-by: Alex Lavaee lavaman131@github.com

Comment on lines +946 to +949
let stopWatching = opened.session.watchRuns?.(next => {
publishRuns(context, next);
if (!chat.busy && !next.active.length && !next.paused.length) void chat.retained?.release();
});

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If the last job finishes while the Planner's reply is being saved, chat.busy is still true, so cleanup gets skipped and never retried. I reproduced the helper and document staying loaded after the job was done.

Let's check again once the reply and any queued replies have finished. Only close the same helper, and only if it still has no work. Please apply this with the cleanup guard below, so a slow close cannot overwrite a newer helper's progress.

Suggested change
let stopWatching = opened.session.watchRuns?.(next => {
publishRuns(context, next);
if (!chat.busy && !next.active.length && !next.paused.length) void chat.retained?.release();
});
let stopWatching = opened.session.watchRuns?.(next => {
if (chat.retained?.session !== opened.session) return;
publishRuns(context, next);
if (next.active.length || next.paused.length) return;
let releaseWhenIdle = async () => {
if (chat.busy) await chat.running;
if (chat.busy || chat.retained?.session !== opened.session) return;
let current = opened.session.runs?.();
if (!current || current.active.length || current.paused.length) return;
await chat.retained.release();
};
void releaseWhenIdle().catch(err =>
console.error("[chat] releasing completed workflow session failed:", err)
);
});

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Applied in c29f101. When the last job ends while a reply, or a queued reply, is still running, cleanup is checked again once it finishes, and it releases only the same helper, only if it still has no active or paused runs. The deferred-cleanup test fails without it.

Assistant-verification: fail-before/pass-after passed: deferred-cleanup test

Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Co-authored-by: Alex Lavaee lavaman131@github.com

Comment thread apps/server/src/chat/service.ts Outdated
Comment on lines +963 to +965
if (chat.agent === opened.session) chat.agent = undefined;
if (!chat.busy) chat.owner = undefined;
publishRuns(context, undefined);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This guard goes with the cleanup change above. Closing an old helper can take long enough for a new one to start. In a regression test, the old cleanup changed the new helper's running card to ‘stopped’. Let's only clear the owner and cards if this is still the helper being closed.

Suggested change
if (chat.agent === opened.session) chat.agent = undefined;
if (!chat.busy) chat.owner = undefined;
publishRuns(context, undefined);
if (chat.agent === opened.session) {
chat.agent = undefined;
if (!chat.busy) chat.owner = undefined;
publishRuns(context, undefined);
}

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in c29f101, generalizing your guard into one rule: every run card belongs to the Planner session that reported it. A report replaces or updates only that session's cards, and releasing a session marks stopped exactly the unended cards it owns, with one "was stopped" message each, saved and announced together. A release never touches another session's cards, and it clears the owner and agent only when the released session is still the current agent and no turn is running. So a slow release of an old helper can no longer overwrite a newer helper's progress, and it still stops its own cards when a newer helper has taken over.

A table test runs four cases through the real invoke flow: no newer helper, a newer helper's run-less turn, a newer helper that reported its own runs, and a run-less turn after a failed save. Each asserts that the old helper's cards end stopped with one message and that the newer helper's cards and owner are untouched.

One known gap remains: if the release's own final save also fails, nothing is announced, and when a newer helper then reports a new run, the old helper's orphaned card is dropped without a stop message. It needs two failures in a row; I can follow up if you'd like it covered here.

Assistant-verification: fail-before/pass-after passed: the four-case release table test (the newer-helper-reported-runs case failed before) and the slow-release tests; bun test 1936 pass, 0 fail on #238 with Atomic 0.9.26-alpha.6
Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Co-authored-by: Alex Lavaee lavaman131@github.com

@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from 24659d0 to c29f101 Compare October 4, 2026 19:39
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from c29f101 to 3d884c7 Compare October 4, 2026 22:35
lavaman131 added a commit that referenced this pull request Oct 4, 2026
…dev/integration

atomic-planner-extensions (#238), with #230 and #215 beneath it, is now
rebased onto main 4430501. main already contains everything else this
branch merged (the earlier main base and #212 and #236), so the result
is the rebased stack's tree.
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from 3d884c7 to 4d1e12f Compare October 5, 2026 05:48
Base automatically changed from atomic-full-planner to main October 5, 2026 05:53
Chopin destroyed every Planner session when its turn ended. A workflow the
atomic Planner starts runs in the background under that session, so it
was orphaned seconds after launch and never progressed. Keep an atomic
Planner session, its owner binding, and the loaded document while the
session still owns live or paused workflow runs, tracked through Atomic's
workflow activity stream; the next turn reuses it so the Planner can still
steer its runs. The session is let go when the runs finish, the owner
binding ends, or the document closes.

Stop Planner now also pauses the session's live runs, resumably, and a
new Resume Planner control (`chat:resume`) resumes them. Chat state
carries the running and paused counts, and the composer shows them.
Run control goes through the session's own workflow commands until the
SDK exposes run-control primitives.

The document's working directory without a checkout moves from the shared
temporary directory to a per-channel directory under Chopin's per-user
state directory, so workflow files survive restarts.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1751 pass, 2 PostgreSQL skips, 0 fail, including new retention, pause, resume, release, and run-classification tests
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: source check passed: DBOS records showed a Planner-launched run with no progress after its turn ended, matching the per-turn session disposal
User-preference: A pause in the Planner pauses its workflows and resume resumes everything; steering goes to the main chat, which decides whether to steer a workflow over intercom
User-preference: Keep long-lived Planner working directories out of the shared tmp
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Document titles are unique per repository, so creating a second document
with an existing title failed in storage and came back as
idempotency-conflict, which reads as a reused key. Report it as
title-taken instead. create_document's output schema also described only
success, so strict MCP clients rejected its idempotency-conflict,
document-unavailable, and validation-issue responses as schema
violations and callers never saw the real outcome; declare them the way
update_document does.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1752 pass, 2 PostgreSQL skips, 0 fail, including a new hosted test for title-taken
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: raw JSON-RPC probe passed: a fresh key with an existing title returned idempotency-conflict before this change, and a new title succeeded
Co-authored-by: Alex Lavaee <lavaman131@github.com>
… prompt

Replace the running/paused count with a card per run, built from Atomic's
workflow lifecycle hooks: the run's name and status (running, waiting on
Decisions, paused, or ended), elapsed time, each stage it reached with
its status and duration, and a link to the Decisions it waits on. Chat
records a line when a run ends. Stop now passes --yes to the pause
command: without it, Atomic asked the host to confirm, which Chopin
routes to Decisions, so the pause waited on a question instead of
pausing. A design-audit specimen covers the running, waiting, paused,
and finished states.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1753 pass, 2 PostgreSQL skips, 0 fail, including run-card folding, paused/waiting/finished status, and the run-ended transcript line
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: Stop Planner on a run waiting in Decisions left "1 workflow running" because the pause waited on its own confirmation
User-preference: Show the workflow's current stages in Chopin, and have pause or quit change what Chat shows immediately
Co-authored-by: Alex Lavaee <lavaman131@github.com>
…se live runs

Stop Planner reported that it paused the workflows, but they kept
running: the /workflow slash command reports its outcome only through
ui.notify, which a headless session drops, so a refused or no-op pause
looked like success. Run pause and resume through the session's own
workflow tool with a real tool context instead, so the outcome returns
as a result and a failure reaches the chat transcript.

The run card's spinner never animated because chat-tool-loader has no
animation in the web app. Live runs and running stages now show a
pulsing dot that stays still under reduced motion.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1754 pass, 2 PostgreSQL skips, 0 fail, including run control through the workflow tool and its failure path
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: Stop Planner posted "paused its workflows" while the run card stayed running and no pause was recorded
User-preference: Live workflow runs in Chopin show a pulsing dot rather than a spinner
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Atomic pauses running and pending stages only, so Stop Planner on a run
whose live stage was waiting on a question changed nothing, while the
transcript said it paused. Read the workflow tool's reported status,
and when a pause changes nothing, quit each live run instead: quitting
keeps the run resumable, withdraws its open question from Decisions,
and asks it again on resume. A run that neither pause nor quit can stop
reports the failure in Chat.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1755 pass, 2 PostgreSQL skips, 0 fail, including pause, quit fallback, and unstoppable-run failure
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: Stop Planner on a run waiting in Decisions left the run card running
User-preference: A pause in the Planner pauses its workflows and resume resumes everything
Co-authored-by: Alex Lavaee <lavaman131@github.com>
…ocked

A run card vanished the moment its last run ended, because Chat kept
cards only while a run was live or paused. Keep ended cards as a one-line
summary with the stage list behind a toggle. A run Atomic ends as
blocked, for example at its review limit, showed as failed; it now shows
as blocked, and the transcript line says it ended blocked.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1755 pass, 2 PostgreSQL skips, 0 fail
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: a run blocked at its review limit disappeared from Chat and its transcript line said "failed after 31 min"
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Atomic's SDK now exposes
session.workflows (bastani-inc/atomic#3377). Stop and Resume now call
session.workflows.pause({ all: true }) and resume each paused run, and a
partial pause reports the runs still active, replacing the workflow tool
called with a hand-built tool context and the quit fallback.

Atomic now counts a run whose stage waits on a question as paused and
leaves the question open, so Stop no longer withdraws it from Decisions.
An answer given while the run is paused is recorded but returned to the
run only after Resume, so the paused run never advances. A run with a
stage waiting on a question now shows as waiting on Decisions even when
no prompt lifecycle event arrives.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1757 pass, 2 PostgreSQL skips, 0 fail, including run control, held answers for paused and child runs, and the waiting count
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
User-preference: Use Atomic's SDK run-control primitives instead of slash commands or tool workarounds
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Chopin counted a workflow root as live only while Atomic reported it
working or blocked. Atomic also reports a live run as idle and quiescent
whenever nothing is counted as executing, for example between steps, so
Chopin released the Planner session mid-run, dropped its card, and left
the run orphaned. A root is now paused when Atomic says so, ended only
once its run's lifecycle reports an end or the root disappears, and live
otherwise.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1757 pass, 2 PostgreSQL skips, 0 fail, including a quiescent root of a live run counted as live
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: on Atomic 0.9.25-alpha.5 the run card vanished eight minutes into grill-me-1 and the run made no further progress
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Atomic withdraws a paused run's open question from Decisions and presents
it again on Resume, so say that instead of describing the question as
staying open.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: agent-browser E2E passed: on Atomic 0.9.25-alpha.5, Stop paused a run waiting on its intent review, Resume presented the same question in Decisions again, and answering it let the run continue to review
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Atomic's workflow guidance tells the model to point people at
/workflow connect, so the Planner told Chopin members to run a command
they have no way to type. The atomic Planner's instructions now say that
members use Chopin in a browser, that slash, terminal, and Atomic CLI
commands are unavailable to them, and that a started workflow is
followed through its run card in Chat, Decisions, and Stop or Resume
Planner, with the Planner offering to check on or steer the run itself.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: session and agent tests (67 pass), including the instruction text for checkout and scratch workspaces and its absence for isolated harnesses
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
User-preference: Chopin members can't run slash commands; the Planner should point them to Chopin's own UI instead of /workflow connect
Co-authored-by: Alex Lavaee <lavaman131@github.com>
Run cards lived only in server memory, so a reload after the last run
ended, or any restart, dropped them. Store them in the document's
sidecar: ended runs stay until the next workflow starts in the document,
and a run that was live when the room went away comes back stopped.

Several runs can be live at once, so Chat now shows one stack of run rows
instead of a card per run. Runs waiting on Decisions come first, then
running, paused, and ended runs, newest first within each; more than three
rows fold behind "more" without hiding a waiting run. Only the first
waiting run, or else the newest live one, shows its stages, and any row
opens on click. Each live or paused row has its own pause or resume
control (chat:pause-run, chat:resume-run) through session.workflows,
recorded in the transcript. Stored stages are capped at the twelve most
recent, with earlier ones counted. A drawn ring replaces the text glyph for
pending and paused work.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1764 pass, 2 PostgreSQL skips, 0 fail, including sidecar round trip, run merging, per-run control, stage cap, and stack ordering and folding
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
User-preference: A finished workflow's card stays until the next workflow run starts, including across reloads
User-preference: Concurrent workflow runs show as one stack ordered by who needs attention, with per-run pause and resume beside Stop all
Co-authored-by: Alex Lavaee <lavaman131@github.com>
A workflow built from ctx.tool steps showed no steps in its run row,
because only stage lifecycle events were folded. Tool steps now appear
in order beside stages, marked with a new Nucleo-style wrench icon and
"tool step" for screen readers; cached counts as done and cancelled as
skipped.

Per-run pause and resume ignored Atomic's outcome, so a pause Atomic
could not apply, such as during a running ctx.tool call, was reported
as done. A noop or cancelled outcome is now a failure with Atomic's
message. Runs of the same workflow are told apart by a short run id in
their rows and control labels.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1767 pass, 2 PostgreSQL skips, 0 fail, including tool steps, the wrench and screen-reader label, and duplicate-name labels
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, type scale, Impeccable (no new findings)
Assistant-verification: agent-browser E2E failed before this change: two concurrent ctx.tool-only runs showed no steps, and pausing one reported "paused" while it kept running
User-preference: Tool stages show a tool icon
Co-authored-by: Alex Lavaee <lavaman131@github.com>
The stack spread a controls object into each row, which the contract
reads as an unreviewed dynamic style; pass the three callbacks
explicitly. The pulse and ring now use the radius and motion tokens
instead of a literal radius, duration, and easing curve. Resume Planner's
icon is the same standalone artwork as Stop Planner's, so it carries the
same reviewed palette exception. WrenchIcon forwards props through
LineIcon like every other line icon, so it joins that reviewed case, and
the line icons and icon catalogue review hashes are renewed for the new
icon and catalogue entries, whose data flows are unchanged.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings); before this change the contract reported the run-card spread, literal radius and motion, the resume icon palette, and the changed icon owners
Assistant-verification: bun test passed: 1902 pass, 2 PostgreSQL skips, 0 fail
Assistant-verification: bun run types passed: all workspaces
Co-authored-by: Alex Lavaee <lavaman131@github.com>
…lpers safely

A run card and its finish message were broadcast before they were saved,
so a failed save left people watching a job finish that a reload then
lost. The update is now built as a candidate outside shared chat state and
saved first: the save stores the candidate transcript and run cards (read
when the commit runs, so messages added meanwhile are kept), and only when
it succeeds are the card and finish message added to the chat and
announced together. While a save is pending, or after it fails, nothing
else that broadcasts chat state or greets a joining client can show the
unsaved card or message. Reports are handled one at a time, so a report
that arrives while an earlier save is pending is computed from the state
after that save settles.

When the save of a report fails, that report is remembered. If the helper
is released afterwards, the release first waits for pending reports and
then publishes the remembered report instead of treating the runs as cut
off, so a transient storage error no longer turns a finished run into a
"was stopped" card; the finished card and its single finish message are
saved and announced together.

Other review fixes from Maggie Appleton on the same code:
- A quick job that ended before the Planner's reply finished now keeps
  its last card and gets its finish message.
- A run first reported as already ended gets one finish message, and a
  run already recorded as ended never gets another.
- When the last job ends while a reply (or a queued reply) is still
  running, cleanup is checked again once it finishes and releases only
  the same helper, and only if it still has no active or paused runs.
- A slow release of an old helper no longer clears the owner or
  overwrites the run cards of a newer helper.

Every run card belongs to the Planner session that reported it, and that
is the only rule for reports and releases. A report from a session
replaces or updates only the cards that session owns and appends its new
ones; live cards of another session still being released are carried
through unchanged, so a newer helper reporting its own runs no longer
drops an older helper's card. Releasing a session replays its own
remembered unsaved report first (a failed final finished report still
publishes as finished), then marks stopped exactly the unended cards it
owns, saving and announcing one "was stopped" message each in a single
save. Cards owned by another session, and the owner and agent of a newer
helper, are never touched; the owner is cleared only when the released
session is still the current agent and no turn is running. If the final
save also fails, nothing is announced and the cards of the released
session are stopped by the next report from any session.

Assistant-model: Claude Sonnet 5.5
Assistant-workflow: goal (run 6091839b-867c-4965-aec9-b583f62d4b1f)
Assistant-verification: table test "releasing an old helper after %s stops exactly its own cards once" (four cases: none, run-less turn, own runs, run-less turn after a failed save) failed before the fix for the own-runs case (run-1 dropped from chat.runs when the newer helper reported run-2; 25 pass, 1 fail in invoke.test.ts) and passes for all four after (26 pass, 0 fail); earlier round-1 to round-5 tests stay green
Assistant-verification: bun test passed: 1918 pass, 2 PostgreSQL skips, 0 fail (private TMPDIR)
Assistant-verification: bun run types passed: all workspaces and e2e
Assistant-verification: bun run ci passed: dprint, oxlint (0 errors, 6 warnings, none in touched files), tokens, design contract, design record, Impeccable
Co-authored-by: Alex Lavaee <lavaman131@github.com>
@lavaman131
lavaman131 force-pushed the atomic-planner-run-control branch from ad5fd51 to 96407eb Compare October 5, 2026 05:56
HARNESS_EXTENSIONS lists extension or package paths that every atomic
Planner session loads, as Atomic's --extension flag would. A package's
extensions, skills and workflows all register, so an operator can give the
Planner workflows without installing them into the agent directory. The
paths are split on the platform path delimiter, must be absolute and exist,
and are refused under any other harness. Background workers never load
them, and the startup line names them.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test passed: 1907 pass, 2 PostgreSQL skips, 0 fail; the new Planner test fails without the adapter change
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint (7 existing warnings, none new), tokens, design contract, design record, Impeccable (no new findings)
Assistant-verification: adapter integration passed: a Planner session with a local package on HARNESS_EXTENSIONS listed that package's workflows through the workflow tool
Co-authored-by: Alex Lavaee <lavaman131@github.com>
…backend

CI timed out the operator-extension test at Bun's 5 s default. Its first
workflow tool call starts Atomic's durable backend, and without Postgres
that falls back to the in-memory backend only after several seconds; with
no Postgres runtime and no Docker, the call alone took 4.7 s locally. The
test is split in two, and both tests that call the workflow tool get 30 s.

Assistant-model: Claude Opus 5.5
Assistant-workflow: inline
Assistant-verification: bun test failed before this change on CI: operator extension paths test timed out after 5000 ms
Assistant-verification: bun test passed: full.test.ts with PATH lacking docker and ATOMIC_POSTGRES_RUNTIME_DIR=/nonexistent (15 pass), and the whole suite (1908 pass, 2 skip, 0 fail)
Assistant-verification: bun run types and bun run ci passed: no new oxlint warnings, no new Impeccable findings
Co-authored-by: Alex Lavaee <lavaman131@github.com>
…e room

A message without @chopin is sent as a room message, and the composer gave
no sign of it, so a member who expected a Planner reply got silence.

The composer now says, below the draft and before anything is sent, "Sends
to the Planner, which will reply" or "Room only. Add @chopin to ask the
Planner". The cue comes from the same chatSendPayload prediction the send
uses, so it follows references and an off Planner (no cue) exactly as the
wire destination will. The @chopin addressing rule and addressed() are
unchanged, and the Send button keeps its name.

Assistant-workflow: goal (run c99248c7-a6cb-4858-95f9-5039c865f0b6)
Assistant-model: Claude Sonnet 5.5
Assistant-verification: bun test passed: apps/web 355 pass, 0 fail, including destinationCue cases for room, planner, reference-masked mention, empty draft and Planner off
Assistant-verification: bun run types and bun run ci passed: dprint, oxlint (7 existing warnings, none new), tokens, design contract, Impeccable (no new findings)
Assistant-verification: playwright e2e passed: bun run e2e e2e/harness.e2e.ts, 2 passed including the composer destination cue test
Co-authored-by: Alex Lavaee <lavaman131@github.com>
0.9.26-alpha.5 relays a workflow stage's ask_user_question to the
session's HostInput (bastani-inc/atomic#3396). On 0.9.25 the relay
missed the request, so a stage's question never reached Decisions and
the stage waited forever. It also restores managed workflow PostgreSQL
after a shutdown (bastani-inc/atomic#3413), which is what left Planner
workflow runs non-durable, so the Planner keeps them durable with
Atomic's own database selection and no extra configuration.

The new full.test.ts case runs a workflow whose stage asks a question
through a real Atomic session: the card appears in Decisions as a
Questionnaire block and its answer reaches the stage. It times out with
no card on 0.9.25. The five builtin packages are unchanged, so
BUILTINS_OFF still turns off each one for the isolated Planner.

Assistant-workflow: inline
Assistant-model: Claude Opus 5.5
Assistant-verification: bun test failed on 0.9.25: the new workflow-stage question case timed out after 32.7s with no card
Assistant-verification: bun test passed: 1911 pass, 2 PostgreSQL skips, 0 fail; full.test.ts 16 pass with the new case
Assistant-verification: bun run types passed: all workspaces
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings)
Assistant-verification: durability check passed: a Planner workflow run with no DBOS_SYSTEM_DATABASE_URL used Atomic's managed PostgreSQL, logged no NON-DURABLY warning, and was kept after the process was killed
Assistant-verification: cross-process /workflow resume from a fresh SDK session failed: timed out on both 0.9.25 and 0.9.26-alpha.5 with persisted checkpoints
Assistant-verification: source check passed: @bastani/atomic 0.9.26-alpha.5 dist/builtin lists intercom, mcp, subagents, web-access, workflows
Co-authored-by: Alex Lavaee <lavaman131@github.com>
0.9.26-alpha.6 lets a new session inspect and resume a workflow run
whose owning session crashed (bastani-inc/atomic#3419). On
0.9.26-alpha.5 only the session that started a run could resume it, so
a Planner session opened after a Chopin restart could never recover the
runs the previous server left behind. The five builtin packages are
unchanged, so BUILTINS_OFF still turns off each one for the isolated
Planner.

Assistant-workflow: inline
Assistant-model: Claude Opus 5.5
Assistant-verification: SDK resume repro passed: after kill -9 and the 2-minute live window, a new session's getRun() reported crashed and resume() completed the run on 0.9.26-alpha.6 (rejected with WorkflowRunOwnershipError on 0.9.26-alpha.5)
Assistant-verification: bun test passed: 1912 pass, 2 PostgreSQL skips, 0 fail
Assistant-verification: bun run types passed: all 9 workspace checks
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings)
Assistant-verification: source check passed: @bastani/atomic 0.9.26-alpha.6 dist/builtin lists intercom, mcp, subagents, web-access, workflows
Co-authored-by: Alex Lavaee <lavaman131@github.com>
0.9.26 is the stable release of the 0.9.26 prereleases already in use: a
workflow stage's ask_user_question reaches Decisions, managed workflow
PostgreSQL comes back after a shutdown, and a new session can resume a
run whose owning session crashed, now with database ownership fencing.
Its one breaking change, McpOAuthCredentialStore taking a server name, is
not used here. @bastani/atomic-natives and @bastani/pi-ai follow to
0.9.26; Atomic's @earendil-works packages stay at 1.0.2. The five builtin
packages are unchanged, so BUILTINS_OFF still turns off each one for the
isolated Planner.

Assistant-workflow: inline
Assistant-model: Claude Opus 5.5
Assistant-verification: bun test passed: 1936 pass, 2 PostgreSQL skips, 0 fail
Assistant-verification: bun run types passed: all 9 workspace checks
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings)
Assistant-verification: source check passed: @bastani/atomic 0.9.26 dist/builtin lists intercom, mcp, subagents, web-access, workflows; no McpOAuthCredentialStore use
Co-authored-by: Alex Lavaee <lavaman131@github.com>
0.9.27-alpha.1 has no breaking changes from 0.9.26. For the Planner it
keeps workflow-stage routing and in-flight workflows reachable across a
failed /reload, delivers background subagent completions to workflow
stages, and stops slow or cancelled OAuth token rotation from dropping a
subscription login. @bastani/atomic-natives and @bastani/pi-ai follow to
0.9.27-alpha.1; Atomic's @earendil-works packages stay at 1.0.2. The five
builtin packages are unchanged, so BUILTINS_OFF still turns off each one
for the isolated Planner.

Assistant-workflow: inline
Assistant-model: Claude Opus 5.5
Assistant-verification: bun test passed: 3665 pass, 3 skips, 0 fail
Assistant-verification: bun run types passed: all 9 workspace checks
Assistant-verification: bun run ci passed: dprint, oxlint, tokens, design contract, design record, Impeccable (no new findings)
Assistant-verification: source check passed: @bastani/atomic 0.9.27-alpha.1 dist/builtin lists intercom, mcp, subagents, web-access, workflows; changelog has no breaking changes since 0.9.26
Co-authored-by: Alex Lavaee <lavaman131@github.com>
@lavaman131
lavaman131 merged commit 0616e52 into main Oct 5, 2026
3 checks passed
@lavaman131
lavaman131 deleted the atomic-planner-run-control branch October 5, 2026 06:13
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants