Repository navigation
[v2] Complete in-process lifecycle and platform reliability work #2525
Description
Activity
- addedsdk-v2Work planned for Copilot SDK v2Work planned for Copilot SDK v2
on Sep 4, 2026 - added a parent issue
on Sep 4, 2026 This issue is a v2 milestone tracking/scoping item — it coordinates assessment and completion criteria for in-process lifecycle and platform reliability work across multiple SDKs, rather than reporting a single bug, requesting a discrete feature, asking a usage question, or flagging a documentation gap. It doesn't fit cleanly into the automated classification categories (bug/enhancement/question/documentation), so it has not been delegated to a handler. A human maintainer should review and triage this tracking issue directly.
Generated by Issue Classification Agent for #2525 · copilot · auto · 8.52 AIC · ⌖ 4.84 AIC · ⊞ 8.7K · ◷
SteveSandersonMS commented
on Sep 4, 2026 ContributorAuthorMore actionsFurther context: napi-oop is no longer used in the runtime, as it no longer depends on a Node child or parent process. We do still have both in-proc and out-of-proc transports and for now we want to continue having CI coverage for both. But any napi-oop-specific concerns can now be discarded.
SteveSandersonMS commented
on Sep 4, 2026 ContributorAuthorMore actionsStatus: SDK-owned work done in a draft PR, one genuine issue escalated upstream
Draft PR: #2531 (kept in draft; CI is green on all 5 in-process SDKs).
What was fixed (SDK-owned)
Root cause: unbounded synchronous native
host_shutdownFFI call, present
in all five in-process SDKs' dispose/close paths (.NET, Node.js, Rust,
Python, Go). A stuck/slow native shutdown — the "SQLite file locking on
Windows" symptom from #1934 — could hang graceful stop and, worse, hang the
documentedforceStop/force_stoprecovery path that exists specifically to
rescue callers from a hang like this. Fixed in all five SDKs: run
host_shutdownon a background thread/task/goroutine, bound the wait
(10s), and defer freeing callback state until the native call actually
completes (never on the timeout path). Added bounded-time regression tests
in every SDK.Also fixed a .NET-specific lifecycle race introduced by bounding
Dispose(): a new client'sStartAsync()could overlap with a previous
client's still-draining, backgrounded shutdown and corrupt shared native
state. Fixed with a semaphore gate serializing native start/shutdown calls.napi-oop is confirmed gone (per earlier maintainer comment on this
issue), so #1983's/#1934's napi-oop-flavored shutdown-crash and SQLite
concerns are moot as root causes. Node.js, Go, Python, and Java's Windows
in-process CI was already enabled and unaffected/still green. Only Rust and
.NET still excluded Windows-in-process from CI going into this work.What was found and could NOT be fixed from the SDK side
Turning on Windows in-process CI for Rust and .NET (to validate the fix
above with real evidence, since the original exclusions were stale) surfaced
a third, deeper, genuine native bug: memory corruption..NET:System.AccessViolationExceptioninConnectionWritecrashing the
whole test process — first during a dispose-adjacent test (looked
lifecycle-related), then again after the lifecycle fix, during a
completely unrelated ordinary RPC call on an already-live connection.Rust:test-inprocess(windows-latest) exits with
STATUS_ACCESS_VIOLATION(SIGSEGV), occurring silently between two
unrelated passing tests, no single deterministic reproducer.
Two independent FFI binding implementations (.NET P/Invoke, Rust
extern "C"/libloading) hit the same fault class calling into the same
nativeconnection_writeexport, on Windows specifically. That's strong
cross-language evidence this is a bug in the shared native runtime
cdylib, not the SDK bindings. This is out of scope for this issue/repo.Filed: github/copilot-agent-runtime#18990 — full stack traces, CI job
links, and the cross-language reasoning.Decision: Windows in-process CI for Rust/.NET has been reverted back to
excluded (matchingmain) in #2531, rather than merging a change that would
make CI reliably red. It should be re-enabled once
github/copilot-agent-runtime#18990 is resolved.Open question / recommendation
Why only .NET and Rust hit this (of the 5 SDKs) is unresolved — flagged as
an open question on the runtime issue rather than asserted. Two candidate
explanations: their E2E suites are much larger so a probabilistic
heap-corruption bug is more likely to surface there by chance, or something
about their specific calling convention/threading/allocator behavior
triggers a latent bug that Node/Go/Python's bindings don't. No maintainer
decision is needed on this issue right now — the recommendation is to review
and merge #2531 for the real, validated SDK-side fixes, and track the
Windows in-process re-enablement as a follow-up once the linked runtime
issue lands.Performance hypothesis from #1934 (25-50% slower shutdown on Windows)
Inconclusive — the native-corruption bug above makes it impossible to get a
meaningful shutdown-latency measurement on Windows in-process right now (a
crashing process has no shutdown time to measure). Revisit once
github/copilot-agent-runtime#18990 is fixed.SteveSandersonMS commented
on Sep 11, 2026 ContributorAuthorMore actionsUpdate after rebasing #2531 onto latest
main(now including Copilot CLI/runtime 1.0.84-4 and #2610's in-process callback-reclamation work):I retried the previously failing Rust/.NET Windows-in-process matrix cells before deciding whether to keep the exclusions. The newer runtime did not clear the blocker:
.NETWindows in-process: all three non-CAPI cells failed again; one reproduced the sameSystem.AccessViolationExceptioninNativeConnectionWrite, and the others crashed/aborted or timed out in the same retry run.RustWindows in-process: still failed when re-enabled. That run also exposed a separate SDK test harness issue in the Rust Windows lifecycle fixture (stdio-only job-object test running under in-process, plus a PID-file readiness race); [v2] Bound in-process FFI host_shutdown and fix .NET lifecycle race across all SDKs #2531 now fixes that harness issue.- Node.js/Go/Python Windows-in-process cells remain green in the normal matrix.
So the recommendation remains: merge #2531 for the SDK-owned bounded-shutdown/lifecycle/test-harness fixes, keep Rust/.NET Windows-in-process excluded for now, and track re-enablement on github/copilot-agent-runtime#18990.
Summary
Complete the remaining in-process runtime lifecycle and platform reliability work required for v2.
These are tracked jobs from #1934, not optional investigation topics. First establish what existing work has already completed, then implement the remainder. If new evidence shows an item no longer applies, document it and obtain maintainer confirmation before closing it without implementation.
Original concerns
napi-oopcleanup race. The race was thought unlikely to affect real applications, but the CI legs still needed to be re-enabled and teardown shown to be robust after removingnapi-oop.forceStop, and possibly not after session disposal. This left locked files on Windows.Required outcomes
napi-oopcleanup work.Implementation preparation
github/copilot-cli, and link the required runtime work.Historical context
Completion
Deliver the required coverage, teardown, resource-release, and shutdown behavior with reliable regression tests. Any item not implemented requires documented contradictory evidence and explicit maintainer agreement.