Describe the bug
In BYOK mode, COPILOT_PROVIDER_WIRE_API (completions or responses) applies to every request in
the session, including sub-agents that run on a different model. Some providers serve each model on
only one of the two APIs. The GitHub Copilot API is one of them: GPT-5 models are served on
/responses and Claude models on /chat/completions. So a custom agent whose model: is in the
other family from the main model can never run. The main agent's first delegation to it fails with
a 400 from the provider.
Copilot CLI already knows the agent's model: subagent.started reports it with
modelSelectionSource: agent_definition_default, and the request goes to the provider with that
model ID. Only the wire API is wrong.
Affected version
GitHub Copilot CLI 1.0.95. Also seen on 1.0.90.
Steps to reproduce the behavior
-
In a git repository, add two custom agents under .github/agents/:
file-summarizer.agent.md with model: claude-haiku-4.5;
quick-checker.agent.md with model: gpt-5.4-mini.
-
GPT main, Claude sub-agent. Start a BYOK session against the Copilot API on the
responses wire API:
COPILOT_PROVIDER_BASE_URL=https://api.githubcopilot.com
COPILOT_PROVIDER_BEARER_TOKEN=<GitHub token with Copilot access>
COPILOT_PROVIDER_HEADERS="Copilot-Integration-Id: copilot-developer-cli"
COPILOT_PROVIDER_WIRE_API=responses
COPILOT_MODEL=gpt-5.4-mini
copilot -p "Delegate to the file-summarizer sub-agent: ask it to summarize README.md. Do not read the file yourself." --allow-all-tools
Output:
✗ File-summarizer (model: claude-haiku-4.5) Summarize README.md
└ 400 model claude-haiku-4.5 does not support Responses API.
-
Claude main, GPT sub-agent. Same settings, but with COPILOT_PROVIDER_WIRE_API=completions
and COPILOT_MODEL=claude-haiku-4.5, delegating to quick-checker:
✗ Quick-checker (model: gpt-5.4-mini) Check Python version in README
└ 400 model "gpt-5.4-mini" is not accessible via the /chat/completions endpoint
-
Control. COPILOT_PROVIDER_WIRE_API=completions with COPILOT_MODEL=claude-sonnet-5,
delegating to file-summarizer (Haiku): the sub-agent runs and returns its summary. Same-family
sub-agents work in both families.
Expected behavior
A sub-agent's requests use a wire API that its own model supports. A custom agent whose model: is
admissible for the provider should run, whatever family the main model is from.
Additional context
Environment.
- Local: Windows 11 Enterprise 10.0.26300, ARM64. Run non-interactively (
copilot -p) from
PowerShell 7.6.
- GitHub Actions: Linux x86_64 runner, Copilot CLI 1.0.90, BYOK through the AWF API proxy (see
below).
- The bug doesn't depend on the OS or terminal: the wire API comes from the session's BYOK
settings.
Log (--log-level debug --log-dir <dir>, GPT main → Haiku sub-agent, prompt bodies omitted).
The sub-agent's request has a Responses API body (input, instructions, include, store),
and the provider rejects it:
[DEBUG] [rust:model_wire] Wire request: {
"model": "claude-haiku-4.5",
(top-level fields: include, input, instructions, max_output_tokens, parallel_tool_calls,
prompt_cache_key, reasoning, store, stream)
[DEBUG] [rust:copilot_runtime::session::run_attempt] Upstream model error response shape (body content omitted) {"status":400,...}
[DEBUG] session.error event received: errorType=query, message=400 model claude-haiku-4.5 does not support Responses API.
Full debug logs are available on request; they contain the session's prompts.
The provider already says which API each model needs. The Copilot API's GET /models returns
supported_endpoints for each model. With the token used above, on 9 Oct:
| Model |
supported_endpoints |
gpt-5.6-luna |
/responses, ws:/responses |
gpt-5.4-mini |
/responses, ws:/responses |
claude-sonnet-5 |
/v1/messages, /chat/completions |
claude-haiku-4.5 |
/chat/completions, /v1/messages |
Possible fixes, in order of preference:
- Choose the wire API per model in BYOK mode. When the provider's
/models lists
supported_endpoints for a model, use a supported wire API for each agent's model, and fall
back to COPILOT_PROVIDER_WIRE_API when the model isn't listed or the field is absent. This
needs the CLI to read the provider's /models, which is close to what copilot-cli#3795 asks for.
- Let the user set the wire API per model or per agent. For example, a mapping from model to
wire API in settings or the environment, or a wire API field on a custom agent definition, used
in preference to the session-wide value.
- At least, fail early and clearly. When an agent's model is known to need another wire API,
say so when the agent is loaded or listed in /subagents, and in the dispatch error, instead of
surfacing the provider's raw 400. (1.0.95 already stops describing an unusable sub-agent as
configured; this would be the same idea for wire API mismatches.)
Why this can't be fixed in a proxy. GitHub Agentic Workflows (gh-aw)
runs Copilot CLI in BYOK mode behind the AWF API proxy,
which forwards to the Copilot API. AWF tried translating between the two APIs (gh-aw-firewall#9558,
fixed by #9559). The requests Copilot CLI sends use features that don't exist on the other API:
include on /responses, and custom tools (tools[custom]). AWF refuses those rather than
silently dropping them, so the sub-agent still fails, now with:
400 Cannot translate Copilot request feature 'include' between Responses and Chat Completions.
Routing model "claude-haiku-4.5" to /chat/completions is incompatible: this request needs
/responses to preserve 'include'.
The client that builds the request is the only place that can pick the right shape.
Why it matters more under model routing. gh-aw can route each run's main model to GPT or
Claude. The CLI's wire API follows the routed pick, so a workflow whose declared sub-agents work on
one pick fails on the other. The workflow author doesn't see it until a run lands on the other
family.
Evidence from GitHub Actions (Copilot CLI 1.0.90, AWF v0.28.49, private githubnext sandbox):
- GPT main (
gpt-5.6-luna) → file-summarizer on claude-haiku-4.5: failed 3 of 3 attempts on
include (run 37813720288).
- Claude main (
claude-sonnet-5) → quick-checker on gpt-5.4-mini: failed 5 of 5 attempts on
tools[custom] (run 37813746913).
- In both runs the same-family sub-agent worked (
small → gpt-5.4-mini under GPT, Haiku under
Sonnet).
- The agents recovered by reporting the missing result or doing the work another way, so the runs
didn't fail, but the delegated work was lost.
Related: gh-aw's own smoke test for sub-agents uses Copilot SDK mode with a gpt-5.3-codex parent
and a claude-haiku-4.5 sub-agent, and it passed as of 6 Oct. We haven't checked how the SDK path
chooses the API, but it suggests the per-model choice already exists outside CLI BYOK mode.
Describe the bug
In BYOK mode,
COPILOT_PROVIDER_WIRE_API(completionsorresponses) applies to every request inthe session, including sub-agents that run on a different model. Some providers serve each model on
only one of the two APIs. The GitHub Copilot API is one of them: GPT-5 models are served on
/responsesand Claude models on/chat/completions. So a custom agent whosemodel:is in theother family from the main model can never run. The main agent's first delegation to it fails with
a 400 from the provider.
Copilot CLI already knows the agent's model:
subagent.startedreports it withmodelSelectionSource: agent_definition_default, and the request goes to the provider with thatmodel ID. Only the wire API is wrong.
Affected version
GitHub Copilot CLI 1.0.95. Also seen on 1.0.90.
Steps to reproduce the behavior
In a git repository, add two custom agents under
.github/agents/:file-summarizer.agent.mdwithmodel: claude-haiku-4.5;quick-checker.agent.mdwithmodel: gpt-5.4-mini.GPT main, Claude sub-agent. Start a BYOK session against the Copilot API on the
responseswire API:Output:
Claude main, GPT sub-agent. Same settings, but with
COPILOT_PROVIDER_WIRE_API=completionsand
COPILOT_MODEL=claude-haiku-4.5, delegating toquick-checker:Control.
COPILOT_PROVIDER_WIRE_API=completionswithCOPILOT_MODEL=claude-sonnet-5,delegating to
file-summarizer(Haiku): the sub-agent runs and returns its summary. Same-familysub-agents work in both families.
Expected behavior
A sub-agent's requests use a wire API that its own model supports. A custom agent whose
model:isadmissible for the provider should run, whatever family the main model is from.
Additional context
Environment.
copilot -p) fromPowerShell 7.6.
below).
settings.
Log (
--log-level debug --log-dir <dir>, GPT main → Haiku sub-agent, prompt bodies omitted).The sub-agent's request has a Responses API body (
input,instructions,include,store),and the provider rejects it:
Full debug logs are available on request; they contain the session's prompts.
The provider already says which API each model needs. The Copilot API's
GET /modelsreturnssupported_endpointsfor each model. With the token used above, on 9 Oct:supported_endpointsgpt-5.6-luna/responses,ws:/responsesgpt-5.4-mini/responses,ws:/responsesclaude-sonnet-5/v1/messages,/chat/completionsclaude-haiku-4.5/chat/completions,/v1/messagesPossible fixes, in order of preference:
/modelslistssupported_endpointsfor a model, use a supported wire API for each agent's model, and fallback to
COPILOT_PROVIDER_WIRE_APIwhen the model isn't listed or the field is absent. Thisneeds the CLI to read the provider's
/models, which is close to what copilot-cli#3795 asks for.wire API in settings or the environment, or a wire API field on a custom agent definition, used
in preference to the session-wide value.
say so when the agent is loaded or listed in
/subagents, and in the dispatch error, instead ofsurfacing the provider's raw 400. (1.0.95 already stops describing an unusable sub-agent as
configured; this would be the same idea for wire API mismatches.)
Why this can't be fixed in a proxy. GitHub Agentic Workflows (gh-aw)
runs Copilot CLI in BYOK mode behind the AWF API proxy,
which forwards to the Copilot API. AWF tried translating between the two APIs (gh-aw-firewall#9558,
fixed by #9559). The requests Copilot CLI sends use features that don't exist on the other API:
includeon/responses, and custom tools (tools[custom]). AWF refuses those rather thansilently dropping them, so the sub-agent still fails, now with:
The client that builds the request is the only place that can pick the right shape.
Why it matters more under model routing. gh-aw can route each run's main model to GPT or
Claude. The CLI's wire API follows the routed pick, so a workflow whose declared sub-agents work on
one pick fails on the other. The workflow author doesn't see it until a run lands on the other
family.
Evidence from GitHub Actions (Copilot CLI 1.0.90, AWF v0.28.49, private githubnext sandbox):
gpt-5.6-luna) →file-summarizeronclaude-haiku-4.5: failed 3 of 3 attempts oninclude(run 37813720288).claude-sonnet-5) →quick-checkerongpt-5.4-mini: failed 5 of 5 attempts ontools[custom](run 37813746913).small→gpt-5.4-miniunder GPT, Haiku underSonnet).
didn't fail, but the delegated work was lost.
Related: gh-aw's own smoke test for sub-agents uses Copilot SDK mode with a
gpt-5.3-codexparentand a
claude-haiku-4.5sub-agent, and it passed as of 6 Oct. We haven't checked how the SDK pathchooses the API, but it suggests the per-model choice already exists outside CLI BYOK mode.