Skip to content

BYOK: sub-agents always use the session's wire API, so a sub-agent on a model from the other family fails with a 400 #5103

Description

@SivaKesava1

Describe the bug

In BYOK mode, COPILOT_PROVIDER_WIRE_API (completions or responses) applies to every request in
the session, including sub-agents that run on a different model. Some providers serve each model on
only one of the two APIs. The GitHub Copilot API is one of them: GPT-5 models are served on
/responses and Claude models on /chat/completions. So a custom agent whose model: is in the
other family from the main model can never run. The main agent's first delegation to it fails with
a 400 from the provider.

Copilot CLI already knows the agent's model: subagent.started reports it with
modelSelectionSource: agent_definition_default, and the request goes to the provider with that
model ID. Only the wire API is wrong.

Affected version

GitHub Copilot CLI 1.0.95. Also seen on 1.0.90.

Steps to reproduce the behavior

  1. In a git repository, add two custom agents under .github/agents/:

    • file-summarizer.agent.md with model: claude-haiku-4.5;
    • quick-checker.agent.md with model: gpt-5.4-mini.
  2. GPT main, Claude sub-agent. Start a BYOK session against the Copilot API on the
    responses wire API:

    COPILOT_PROVIDER_BASE_URL=https://api.githubcopilot.com
    COPILOT_PROVIDER_BEARER_TOKEN=<GitHub token with Copilot access>
    COPILOT_PROVIDER_HEADERS="Copilot-Integration-Id: copilot-developer-cli"
    COPILOT_PROVIDER_WIRE_API=responses
    COPILOT_MODEL=gpt-5.4-mini
    copilot -p "Delegate to the file-summarizer sub-agent: ask it to summarize README.md. Do not read the file yourself." --allow-all-tools
    

    Output:

    ✗ File-summarizer (model: claude-haiku-4.5) Summarize README.md
      └ 400 model claude-haiku-4.5 does not support Responses API.
    
  3. Claude main, GPT sub-agent. Same settings, but with COPILOT_PROVIDER_WIRE_API=completions
    and COPILOT_MODEL=claude-haiku-4.5, delegating to quick-checker:

    ✗ Quick-checker (model: gpt-5.4-mini) Check Python version in README
      └ 400 model "gpt-5.4-mini" is not accessible via the /chat/completions endpoint
    
  4. Control. COPILOT_PROVIDER_WIRE_API=completions with COPILOT_MODEL=claude-sonnet-5,
    delegating to file-summarizer (Haiku): the sub-agent runs and returns its summary. Same-family
    sub-agents work in both families.

Expected behavior

A sub-agent's requests use a wire API that its own model supports. A custom agent whose model: is
admissible for the provider should run, whatever family the main model is from.

Additional context

Environment.

  • Local: Windows 11 Enterprise 10.0.26300, ARM64. Run non-interactively (copilot -p) from
    PowerShell 7.6.
  • GitHub Actions: Linux x86_64 runner, Copilot CLI 1.0.90, BYOK through the AWF API proxy (see
    below).
  • The bug doesn't depend on the OS or terminal: the wire API comes from the session's BYOK
    settings.

Log (--log-level debug --log-dir <dir>, GPT main → Haiku sub-agent, prompt bodies omitted).
The sub-agent's request has a Responses API body (input, instructions, include, store),
and the provider rejects it:

[DEBUG] [rust:model_wire] Wire request: {
  "model": "claude-haiku-4.5",
  (top-level fields: include, input, instructions, max_output_tokens, parallel_tool_calls,
   prompt_cache_key, reasoning, store, stream)
[DEBUG] [rust:copilot_runtime::session::run_attempt] Upstream model error response shape (body content omitted) {"status":400,...}
[DEBUG] session.error event received: errorType=query, message=400 model claude-haiku-4.5 does not support Responses API.

Full debug logs are available on request; they contain the session's prompts.

The provider already says which API each model needs. The Copilot API's GET /models returns
supported_endpoints for each model. With the token used above, on 9 Oct:

Model supported_endpoints
gpt-5.6-luna /responses, ws:/responses
gpt-5.4-mini /responses, ws:/responses
claude-sonnet-5 /v1/messages, /chat/completions
claude-haiku-4.5 /chat/completions, /v1/messages

Possible fixes, in order of preference:

  1. Choose the wire API per model in BYOK mode. When the provider's /models lists
    supported_endpoints for a model, use a supported wire API for each agent's model, and fall
    back to COPILOT_PROVIDER_WIRE_API when the model isn't listed or the field is absent. This
    needs the CLI to read the provider's /models, which is close to what copilot-cli#3795 asks for.
  2. Let the user set the wire API per model or per agent. For example, a mapping from model to
    wire API in settings or the environment, or a wire API field on a custom agent definition, used
    in preference to the session-wide value.
  3. At least, fail early and clearly. When an agent's model is known to need another wire API,
    say so when the agent is loaded or listed in /subagents, and in the dispatch error, instead of
    surfacing the provider's raw 400. (1.0.95 already stops describing an unusable sub-agent as
    configured; this would be the same idea for wire API mismatches.)

Why this can't be fixed in a proxy. GitHub Agentic Workflows (gh-aw)
runs Copilot CLI in BYOK mode behind the AWF API proxy,
which forwards to the Copilot API. AWF tried translating between the two APIs (gh-aw-firewall#9558,
fixed by #9559). The requests Copilot CLI sends use features that don't exist on the other API:
include on /responses, and custom tools (tools[custom]). AWF refuses those rather than
silently dropping them, so the sub-agent still fails, now with:

400 Cannot translate Copilot request feature 'include' between Responses and Chat Completions.
Routing model "claude-haiku-4.5" to /chat/completions is incompatible: this request needs
/responses to preserve 'include'.

The client that builds the request is the only place that can pick the right shape.

Why it matters more under model routing. gh-aw can route each run's main model to GPT or
Claude. The CLI's wire API follows the routed pick, so a workflow whose declared sub-agents work on
one pick fails on the other. The workflow author doesn't see it until a run lands on the other
family.

Evidence from GitHub Actions (Copilot CLI 1.0.90, AWF v0.28.49, private githubnext sandbox):

  • GPT main (gpt-5.6-luna) → file-summarizer on claude-haiku-4.5: failed 3 of 3 attempts on
    include (run 37813720288).
  • Claude main (claude-sonnet-5) → quick-checker on gpt-5.4-mini: failed 5 of 5 attempts on
    tools[custom] (run 37813746913).
  • In both runs the same-family sub-agent worked (small → gpt-5.4-mini under GPT, Haiku under
    Sonnet).
  • The agents recovered by reporting the missing result or doing the work another way, so the runs
    didn't fail, but the delegated work was lost.

Related: gh-aw's own smoke test for sub-agents uses Copilot SDK mode with a gpt-5.3-codex parent
and a claude-haiku-4.5 sub-agent, and it passed as of 6 Oct. We haven't checked how the SDK path
chooses the API, but it suggests the per-model choice already exists outside CLI BYOK mode.

Activity

  1. added theissue type on Oct 9, 2026
  2. added
    area:agentsSub-agents, fleet, autopilot, plan mode, background agents, and custom agents
    area:modelsModel selection, availability, switching, rate limits, and model-specific behavior
    and removed on Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:agentsSub-agents, fleet, autopilot, plan mode, background agents, and custom agentsarea:modelsModel selection, availability, switching, rate limits, and model-specific behavior

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions