Skip to content

Hosted idle hibernation interrupts active tools while local shell keeps running #183

Description

@iamnbutler

A hosted channel with active but quiet workspace work can hibernate after ordinary client sockets close. The cell reconstructs its runtime and loses the in-memory pending workspace call; pi reports the tool as interrupted and continues the agent while the original local shell is still running. This is a reproduced deployed Cloudflare defect from pilot #176, not a local-runtime inference.

Fixed — October 7

PR #186, reviewed head 4d534550b0a36d9c8272113018c71c0b9b199830, merged as 80c67ca on October 7 at 13:37:52 UTC after GitHub CI passed. It adds the service-only busy residency timer and documents it. The channel storage and workspace protocol are unchanged. This issue closed automatically on merge.

The deployed candidate passed the quiet-tool, delayed Kill, ordinary Stop/Kill, subsequent idle hibernation and genuine-reset checks below. Full pilot evidence distinguishes plain-deploy continuity from actual ctx.abort() resets. Final deployed version b58efd2d-579e-4c70-a9aa-c326bca1e086 uses the original service entry point; the temporary reset wrapper is removed and the same channel/history remain. These deployed results were recorded before merge and remain tied to that pilot version.

Reproduction and evidence

  • Source revision: ad3f7313ea18f2fd9c5f9e3fb2ca9c5b1df67861, deployed as the dedicated ace-channel-pilot service on real Cloudflare. Real Sonnet model and the actual local workspace adapter.
  • On October 7, a detached run started one silent 75-second bash tool at 11:35:03 UTC, under runtime generation 0690e4c9…, local PID 15079. The tool call id was toolu_01RXmwrpZakAKS869PQuHNrG.
  • All ordinary client sockets were closed. The workspace socket stayed connected throughout; there was no service deployment or intentional workspace disconnect.
  • At 11:35:38 UTC, a new runtime generation 2b08a256… started. Pi recorded the original call as “Tool bash was interrupted and may have partially run”, while its original PID and sleep child remained alive on the workspace.
  • The agent avoided retrying the original command and instead checked/waited. A later wait tool was also reported interrupted after another runtime recovery.
  • The original shell eventually wrote its end marker at 11:36:18 UTC, confirming that execution continued after the channel had already recorded interruption and advanced.

A second pre-fix case (T5) confirms a specific Kill consequence after that loss of pending-call state:

  • A silent 120-second shell started at 11:37:09 UTC, generation fdc6b613:6, PID 20265 with sleep child 20268.
  • Pi reported interruption roughly 30 seconds later. The model followed its no-retry instruction and finished with “T5 interrupted” at 11:37:38 UTC; pi therefore considered the chat idle while the original shell was still running.
  • Hosted Kill at 11:38:12 UTC returned successfully, but the original PID 20265 and sleep child 20268 remained alive afterward. Kill could no longer reach the old workspace call orphaned by hibernation.

This proves Kill after hibernation cannot cancel that orphaned call. It does not establish that ordinary Kill fails while an active workspace call remains attached to its runtime.

Bounded process, marker, replay and runtime-generation evidence is retained by the pilot operator. The public scenario summary records the reproduction and regressions without machine-private paths, credentials or raw transcripts.

Required behavior and validation

These checks passed on the deployed candidate before the validated fix merged in #186.

  • Keep the runtime resident while any chat is busy, including a quiet pending workspace tool. PR fix(channel): keep hosted runs resident while busy #186 uses the existing busy callback and a bounded re-arming timer; the idle transition clears it.
  • Repeat the detached quiet-shell reproduction: the final timer version completed a 75-second silent shell once in one generation with no clients, interrupted result, retry or workspace drop (T4 final).
  • Verify ordinary Stop and Kill remove the shell and child within three seconds with no end marker (T8/T9). Repeat the former orphaning window: a silent 120-second shell with no clients for 56 seconds received Kill from its original generation; both processes exited within three seconds and no end marker appeared after its original deadline (T11).
  • Allow idle hibernation again after the run ends: a later tool used a new generation and succeeded while the workspace socket survived (T7). Channel history and usage comparisons remain consistent across deployments.
  • Preserve durable alarm/pi recovery for genuine resets: temporary pilot-only ctx.abort() during a shell dropped/reconnected the workspace, removed the old processes and reported interruption without unsafe replay (T13). A model-stream reset retained the stopped partial reply and produced a complete continuation (T14). A plain deploy did not demonstrate a reset; the temporary reset wrapper was removed afterward.
  • Review and merge fix(channel): keep hosted runs resident while busy #186; the validated fix landed as 80c67ca.

Keep the fix focused on the service's busy/residency lifecycle. A broader workspace protocol rewrite is not established as necessary by this finding. Preserve runtime-neutral channel code, pi's durable state and existing channels; no store reset.

First blocker found in deployed pilot #176, fixed in #186. Hosted-channel meta #13 and dogfooding #5 retain the finding; the pilot remains open for its separate physical-machine checks.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions