Skip to content

[v2] Complete in-process lifecycle and platform reliability work #2525

Description

@SteveSandersonMS

Summary

Complete the remaining in-process runtime lifecycle and platform reliability work required for v2.

These are tracked jobs from #1934, not optional investigation topics. First establish what existing work has already completed, then implement the remainder. If new evidence shows an item no longer applies, document it and obtain maintainer confirmation before closing it without implementation.

Original concerns

  • Windows in-process CI legs for .NET and Rust were disabled because of a napi-oop cleanup race. The race was thought unlikely to affect real applications, but the CI legs still needed to be re-enabled and teardown shown to be robust after removing napi-oop.
  • The underlying runtime did not close SQLite connections after forceStop, and possibly not after session disposal. This left locked files on Windows.
  • Graceful SQLite shutdown was suspected of being slow. This was the leading hypothesis in [Tracking] In-process (FFI) items to be cleaned up #1934 for in-process CI legs taking approximately 25–50% longer than their out-of-process equivalents.

Required outcomes

  • The intended Windows in-process CI coverage is enabled and reliable across SDKs.
  • Callback and teardown behavior remains robust after the napi-oop cleanup work.
  • Force-stop and session disposal release SQLite connections and do not leave locked files on Windows.
  • Graceful shutdown does not have unexplained material performance regressions relative to out-of-process operation.
  • Any intentionally unsupported platform or test combination is explicitly justified and documented.

Implementation preparation

Historical context

Completion

Deliver the required coverage, teardown, resource-release, and shutdown behavior with reliable regression tests. Any item not implemented requires documented contradictory evidence and explicit maintainer agreement.

Activity

  1. github-actions commented on Sep 4, 2026

    @github-actions
    Contributor

    This issue is a v2 milestone tracking/scoping item — it coordinates assessment and completion criteria for in-process lifecycle and platform reliability work across multiple SDKs, rather than reporting a single bug, requesting a discrete feature, asking a usage question, or flagging a documentation gap. It doesn't fit cleanly into the automated classification categories (bug/enhancement/question/documentation), so it has not been delegated to a handler. A human maintainer should review and triage this tracking issue directly.

    Generated by Issue Classification Agent for #2525 · copilot · auto · 8.52 AIC · ⌖ 4.84 AIC · ⊞ 8.7K · ◷

  2. SteveSandersonMS commented on Sep 4, 2026

    @SteveSandersonMS
    ContributorAuthor

    Further context: napi-oop is no longer used in the runtime, as it no longer depends on a Node child or parent process. We do still have both in-proc and out-of-proc transports and for now we want to continue having CI coverage for both. But any napi-oop-specific concerns can now be discarded.

  3. SteveSandersonMS commented on Sep 4, 2026

    @SteveSandersonMS
    ContributorAuthor

    Status: SDK-owned work done in a draft PR, one genuine issue escalated upstream

    Draft PR: #2531 (kept in draft; CI is green on all 5 in-process SDKs).

    What was fixed (SDK-owned)

    Root cause: unbounded synchronous native host_shutdown FFI call, present
    in all five in-process SDKs' dispose/close paths (.NET, Node.js, Rust,
    Python, Go). A stuck/slow native shutdown — the "SQLite file locking on
    Windows" symptom from #1934 — could hang graceful stop and, worse, hang the
    documented forceStop/force_stop recovery path that exists specifically to
    rescue callers from a hang like this. Fixed in all five SDKs: run
    host_shutdown on a background thread/task/goroutine, bound the wait
    (10s), and defer freeing callback state until the native call actually
    completes (never on the timeout path). Added bounded-time regression tests
    in every SDK.

    Also fixed a .NET-specific lifecycle race introduced by bounding
    Dispose(): a new client's StartAsync() could overlap with a previous
    client's still-draining, backgrounded shutdown and corrupt shared native
    state. Fixed with a semaphore gate serializing native start/shutdown calls.

    napi-oop is confirmed gone (per earlier maintainer comment on this
    issue), so #1983's/#1934's napi-oop-flavored shutdown-crash and SQLite
    concerns are moot as root causes. Node.js, Go, Python, and Java's Windows
    in-process CI was already enabled and unaffected/still green. Only Rust and
    .NET still excluded Windows-in-process from CI going into this work.

    What was found and could NOT be fixed from the SDK side

    Turning on Windows in-process CI for Rust and .NET (to validate the fix
    above with real evidence, since the original exclusions were stale) surfaced
    a third, deeper, genuine native bug: memory corruption.

    • .NET: System.AccessViolationException in ConnectionWrite crashing the
      whole test process — first during a dispose-adjacent test (looked
      lifecycle-related), then again after the lifecycle fix, during a
      completely unrelated ordinary RPC call on an already-live connection.
    • Rust: test-inprocess (windows-latest) exits with
      STATUS_ACCESS_VIOLATION (SIGSEGV), occurring silently between two
      unrelated passing tests, no single deterministic reproducer.

    Two independent FFI binding implementations (.NET P/Invoke, Rust
    extern "C"/libloading) hit the same fault class calling into the same
    native connection_write export, on Windows specifically. That's strong
    cross-language evidence this is a bug in the shared native runtime
    cdylib
    , not the SDK bindings. This is out of scope for this issue/repo.

    Filed: github/copilot-agent-runtime#18990 — full stack traces, CI job
    links, and the cross-language reasoning.

    Decision: Windows in-process CI for Rust/.NET has been reverted back to
    excluded (matching main) in #2531, rather than merging a change that would
    make CI reliably red. It should be re-enabled once
    github/copilot-agent-runtime#18990 is resolved.

    Open question / recommendation

    Why only .NET and Rust hit this (of the 5 SDKs) is unresolved — flagged as
    an open question on the runtime issue rather than asserted. Two candidate
    explanations: their E2E suites are much larger so a probabilistic
    heap-corruption bug is more likely to surface there by chance, or something
    about their specific calling convention/threading/allocator behavior
    triggers a latent bug that Node/Go/Python's bindings don't. No maintainer
    decision is needed on this issue right now — the recommendation is to review
    and merge #2531 for the real, validated SDK-side fixes, and track the
    Windows in-process re-enablement as a follow-up once the linked runtime
    issue lands.

    Performance hypothesis from #1934 (25-50% slower shutdown on Windows)

    Inconclusive — the native-corruption bug above makes it impossible to get a
    meaningful shutdown-latency measurement on Windows in-process right now (a
    crashing process has no shutdown time to measure). Revisit once
    github/copilot-agent-runtime#18990 is fixed.

  4. SteveSandersonMS commented on Sep 11, 2026

    @SteveSandersonMS
    ContributorAuthor

    Update after rebasing #2531 onto latest main (now including Copilot CLI/runtime 1.0.84-4 and #2610's in-process callback-reclamation work):

    I retried the previously failing Rust/.NET Windows-in-process matrix cells before deciding whether to keep the exclusions. The newer runtime did not clear the blocker:

    • .NET Windows in-process: all three non-CAPI cells failed again; one reproduced the same System.AccessViolationException in NativeConnectionWrite, and the others crashed/aborted or timed out in the same retry run.
    • Rust Windows in-process: still failed when re-enabled. That run also exposed a separate SDK test harness issue in the Rust Windows lifecycle fixture (stdio-only job-object test running under in-process, plus a PID-file readiness race); [v2] Bound in-process FFI host_shutdown and fix .NET lifecycle race across all SDKs #2531 now fixes that harness issue.
    • Node.js/Go/Python Windows-in-process cells remain green in the normal matrix.

    So the recommendation remains: merge #2531 for the SDK-owned bounded-shutdown/lifecycle/test-harness fixes, keep Rust/.NET Windows-in-process excluded for now, and track re-enablement on github/copilot-agent-runtime#18990.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    sdk-v2Work planned for Copilot SDK v2

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions