Skip to content

Allow main agents to resume subagent conversations - #614

Open
itkonen wants to merge 8 commits into
editor-code-assistant:masterfrom
itkonen:feat/resumable-subagents
Open

itkonen wants to merge 8 commits into
editor-code-assistant:masterfrom
itkonen:feat/resumable-subagents

Conversation

@itkonen

@itkonen itkonen commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

This PR allows main agents to reuse subagents and their accumulated context to speed up orchestration. Similar features exist in harnesses such as OpenCode, Codex, and Claude Code. Here, this is implemented by allowing the spawn_agent tool to accept the chat_id of an earlier subagent. When a chat_id is provided, the subagent continues its conversation from where it left off.

This supports several communication patterns. For example, a subagent can request more details from the main agent and return. The main agent can then answer through the task argument, using chat_id to continue that same subagent conversation.

AI summary

Implementation

  • Adds optional chat_id to spawn_agent. Calls without it continue to create fresh subagents. agent and task remain required.
  • Reuses the existing chat-prompt machinery and retained child history, including tool calls and results. Existing context compaction still applies.
  • Preserves the child’s original resolved model and variant. Continuation requires the same agent and does not accept model or variant overrides.
  • Returns the reusable ID at the beginning of the tool result so normal output truncation preserves it.
  • Resets the step budget for each invocation and extracts its result only from newly appended history, preventing an earlier answer from being returned as the new result.

Ownership and lifecycle safeguards

Continuation is limited to the same parent chat and server session. Invalid, foreign, or unauthorized IDs are rejected without silently creating a fresh child. Changes to configuration, workspace, or trust also require a fresh subagent. History replay checks child ownership before expanding referenced conversations.

Session-local bookkeeping prevents overlapping invocations and waits for prompt workers and tool cleanup to finish. Cancelled or provider-failed runs can be continued once their previous work has settled; active tools or unfinished cleanup still block continuation.

The patch also makes targeted corrections to existing cancellation handling:

  • Tool completion is signalled after post-tool and status hooks finish.
  • Stopping no longer bypasses the wait for dispatched tools and their cleanup.
  • Tool-state callbacks read current state rather than a captured snapshot.

There is no new messaging protocol, scheduler, or dedicated parent-question tool. Fork-aware continuation and persistence across server restarts are outside scope.

Validation

Regression coverage includes retained conversation and tool history, fresh-child behavior, invalid and unauthorized IDs, parameter ordering, model preservation, step-limit continuation, truncated results, transient provider failures, cancellation races, and rejection while old work remains active.

  • Full unit suite: 976 tests, 5,974 assertions, zero failures, using an explicit English/US JVM locale because existing numeric-format tests assume decimal points.
  • Existing subagent integration tests: 2 tests, 13 assertions, zero failures or errors.
  • Changed-file lint, editor diagnostics, and whitespace checks: clean.

🤖 Generated with ECA (openai/gpt-6-astra - xhigh)

Let main agents reuse a child's accumulated context through spawn_agent's
optional chat_id instead of recreating the conversation for every task.
Preserve the original model settings and enforce same-parent ownership,
session-local authorization, and non-overlapping invocations.

Allow cancelled and provider-failed runs to continue after their work has
settled, and wait for dispatched tools and cleanup hooks before reuse.
Cover retained history, validation, retry, and cancellation races with
regression tests and document the continuation interface.

🤖 Generated with [ECA](https://eca.dev) (openai/gpt-6-astra - xhigh)

Co-Authored-By: eca-agent <git@eca.dev>
@itkonen
itkonen marked this pull request as ready for review September 20, 2026 12:39
itkonen and others added 3 commits September 22, 2026 20:58
Account for plugin-update and the updated plugin-uninstall argument metadata. Remove the unrelated eca-info description assertion from the previous fix.
@zikajk

zikajk commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Personally, I think we need this functionality, so thank you for working on it @itkonen.

I do wonder, though, whether a more general orchestration model like the one used by Claude Code or Codex might be better in the long run. Chats and subagents could share the same underlying session model, with agents able to communicate with one another.

(when using Opus, you can sometimes see the main agent give subagents instructions that assume they have access to the main chat's context).

I don’t see any bugs, but some of the code Astra generated feels rather awkward for Clojure. Maybe it would be worth running it through an Anthropic model, but I’ll leave the decision to @ericdallo.

@itkonen

itkonen commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

Thank you for the comment @zikajk!

I also briefly considered more elaborate communication tools, but I came to the conclusion that this simple solution would already allow for quite many communication patterns – although it might require a bit more from the agent instructions, as agents might not figure it out by themselves.

And I must apologize for the Astra code. I've found it almost impossible to force Astra to keep it simple and focus on the main quest, and not create to a plethora of validators and tests. Yet it is quite clever in figuring out solutions to complex problems.

@zikajk

zikajk commented Sep 24, 2026

Copy link
Copy Markdown
Member

Totally agree, Astra is smart but not the best worker :-).

itkonen and others added 2 commits September 29, 2026 21:22
- Remove unused and redundant functions related to subagent summary and max-steps
- Consolidate and clarify resume validation logic for subagents
- Update tests to reflect new resume behavior and remove obsolete test cases
- Ensure assistant text extraction returns nil when no assistant messages
- Clean up test code and assertions for subagent worker and resume scenarios
Bring in upstream fixes through edca860 while preserving the
resumable-subagent cleanup and Unreleased changelog entry.

🤖 Generated with [ECA](https://eca.dev) (openai/gpt-6-astra - xhigh)

Co-Authored-By: eca-agent <git@eca.dev>
Comment thread resources/prompts/tools/spawn_agent.md Outdated
- 'task': Provide a highly detailed prompt. Explicitly state whether it should write/edit code or just research, how to verify its work, and exactly what specific information it must return to you.
- 'activity': Must be a concise 3-4 word label for the UI (e.g., "exploring codebase", "refactoring module").
- 'activity': Optional concise 3-4 word label for the UI (e.g., "exploring codebase", "refactoring module").
- 'chat_id': Optional returned ID to continue a conversation in the same parent chat and server session. Reuse its 'agent', supply a new 'task', and omit 'model' and 'variant'.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if this is already enough for agents to use this and continue how your tests went?

@ericdallo ericdallo left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good overall, can you fix conflicts so I take a final look pls?

@zikajk

zikajk commented Oct 8, 2026

Copy link
Copy Markdown
Member

@ericdallo @itkonen I will do it (as a new commit), I hope you don't mind.

@zikajk
zikajk force-pushed the feat/resumable-subagents branch from 1b9383a to 688816c Compare October 9, 2026 13:25
Resolve the conflict with the subagent timeoutSeconds and final summary
turn (editor-code-assistant#625):

- Pass the subagent token in the shared prompt params, so the summary
  turn is not rejected by the managed-subagent prompt guard.
- Build the Halted, Timed out and parent-stop texts only from this
  invocation's messages, so a continued subagent never returns an answer
  from an earlier run.
- Expect the extra summary-turn request in the max-steps resume test.
@zikajk
zikajk force-pushed the feat/resumable-subagents branch from 688816c to c4ede97 Compare October 9, 2026 13:31
Follow-up to the original commits, on top of the master merge (which
resolves the conflict with editor-code-assistant#625). Changes after the original work:

Fixes:
- Clear `:summary-requested?` when a subagent is continued and tell it
  that tool calls are allowed again; after a editor-code-assistant#625 summary turn it could
  never call tools.
- Wait for a stopped turn to settle before the summary turn and again
  after it (one shared deadline), so late tool results do not mix in and
  the subagent can be continued at once. Time out only a running turn.
- Replay only the part of a continued subagent chat that each call ran:
  the call records a `[start, end)` range (`subagentMessageRange`,
  `[0, 0]` when it did not run). Old histories still replay everything.
- After a parent stop, wait up to 5 s (polling every 50 ms) for the
  subagent to settle, so the cancelled tool result is replayed and the
  subagent can be continued right away.
- A direct prompt to a live subagent chat returns the normal error
  result instead of throwing. A run that was only interrupted is
  reported as stopped, not as a generic failure. No `variant: null` in
  the final tool-call details.

Behaviour changes to the original design:
- Drop the config, workspace and trust check on continue. The config
  hash changes whenever a `${cmd:...}` value, a config file or a plugin
  changes, which blocked continuing after the user answers a question.
  A continued subagent gets the parent's current trust like a new one,
  and its agent is authorized again against the current config.
- `maxSteps` and `timeoutSeconds` come from the current agent config,
  with a fresh budget for each run.
- Allow `model` and `variant` on continue. Without them the subagent
  keeps its model and variant; a new model gets the agent's configured
  variant, and the variant is validated against the final model.
- An empty or null `chat_id` spawns a new subagent, like other empty
  optional arguments; revert the `minLength` case in the generic helper.
- One continue error in terms the parent model knows, saying how to
  recover (omit `chat_id` to spawn a new subagent).

Simplifications:
- Session-only run data holds just the token, worker count and
  interrupted flag; parent, agent, model and variant live once, on the
  subagent chat. Fewer `:interrupted?` writes; the worker counter is a
  plain try/finally.
- spawn_agent is split into small functions: pure `admit-new`,
  `admit-resume` and `release` db functions used in `swap!`, a run map
  built from the handler context, and one `prompt-subagent!` for the
  task and the summary turn.
- The tool-call cleanup signal stays in the state table as
  `:finally-actions` instead of a hard-coded event set.
- The chat side only knows a generic managed chat (`:managed-chats`:
  owner token, worker count, interrupted flag; `:owner-token` in prompt
  params). Subagent logic, including the replay slice, is in agent.clj.

Docs (agents, protocol, tool prompt, CHANGELOG) and tests updated.
The tool prompt explains recovery when a subagent is unavailable,
instead of asking the model to identify a server session. The restart
limit remains in the human documentation.
@zikajk
zikajk force-pushed the feat/resumable-subagents branch from c4ede97 to 40de443 Compare October 9, 2026 13:55
@zikajk

zikajk commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

@itkonen I merged master and put all my changes in one commit on top of yours. Its message lists every change. The main ones: I dropped the config/trust check (it blocked continuing after config changes), and model/variant are now allowed on continue.

Could you test it with a real model? Mainly: a subagent returns a request for clarification, the parent asks the user, then continues the same subagent with the answer. Also test continuation after maxSteps/timeout and after a stop.

One open problem: after an ECA restart, the parent still sees the old chat_id in its history, but continuing it is rejected. Could you check how big a problem this is in practice, e.g. does it confuse the model, or does it just spawn a new subagent? I still plan to do a follow-up though. It doesn't look like a big change (although I can see us throwing away a lot of code if we start working on Agent Swarm)

@ericdallo You might want to take a look, there are some bigger changes.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants