diff --git a/docs/features/streaming-events.md b/docs/features/streaming-events.md
index b20339042a..c6bdd906a0 100644
--- a/docs/features/streaming-events.md
+++ b/docs/features/streaming-events.md
@@ -1017,7 +1017,7 @@ This table lists key `data` payload fields. Common envelope fields are documente
| `assistant.message` | | Assistant | `messageId`, `content`, `toolRequests?`, `outputTokens?`, `phase?` |
| `assistant.message_delta` | ✅ | Assistant | `messageId`, `deltaContent` |
| `assistant.turn_end` | | Assistant | `turnId` |
-| `assistant.usage` | ✅ | Assistant | `model`, `apiEndpoint?`, `inputTokens?`, `outputTokens?`, `cost?`, `duration?` |
+| `assistant.usage` | ✅ | Assistant | `model`, `apiEndpoint?`, `inputTokens?`, `outputTokens?`, `cacheReadTokens?`, `cacheWriteTokens?`, `cost?`, `duration?` |
| `tool.user_requested` | | Tool | `toolCallId`, `toolName`, `arguments?` |
| `tool.execution_start` | | Tool | `toolCallId`, `toolName`, `arguments?`, `mcpServerName?` |
| `tool.execution_partial_result` | ✅ | Tool | `toolCallId`, `partialOutput` |
diff --git a/docs/features/usage-and-billing.md b/docs/features/usage-and-billing.md
index c412cb35c2..4eb63e78c9 100644
--- a/docs/features/usage-and-billing.md
+++ b/docs/features/usage-and-billing.md
@@ -38,21 +38,28 @@ Usage totals belong to a session, not the active account. In the CLI, switching
> [!NOTE]
> `session.quota`, `session.usage.getMetrics`, `session.metadata.contextInfo`, and `session.metadata.recomputeContextTokens` are marked experimental in the generated RPC surface. In .NET they raise the `GHCP001` experimental diagnostic, which you suppress with `#pragma warning disable GHCP001` or a project-level `GHCP001`. Pin both the SDK and the Copilot CLI runtime if your application depends on them.
-The field tables below list only the fields used in the examples on this page. The complete, always-current field reference is the generated SDK types plus [Streaming events](./streaming-events.md), which is regenerated from the CLI schema on every dependency bump. Treat those as the source of truth and this page as a task-oriented guide.
+The field tables below highlight the usage data covered on this page. The complete, always-current field reference is the generated SDK types plus [Streaming events](./streaming-events.md), which is regenerated from the CLI schema on every dependency bump. Treat those as the source of truth and this page as a task-oriented guide.
## Per-call token counts
The `assistant.usage` event is emitted once for every model API call in a turn (including calls made by sub-agents). It carries the token counts and the billing multiplier for that single call.
-The example below uses these fields. See [Streaming events](./streaming-events.md#assistantusage) for the full list, including cache, reasoning, latency, and tracing fields.
+The example below uses these fields. See [Streaming events](./streaming-events.md#assistantusage) for the full list, including reasoning, latency, and tracing fields.
| Field | Type | Description |
|---|---|---|
| `model` | `string` | Model identifier for this call |
| `inputTokens` | `number` | Input tokens consumed |
| `outputTokens` | `number` | Output tokens produced |
+| `cacheReadTokens` | `number` (optional) | Tokens read from the prompt cache for this call |
+| `cacheWriteTokens` | `number` (optional) | Tokens written to the prompt cache for this call |
| `cost` | `number` | Premium request multiplier applied to this call |
+An absent cache count means no value was reported. The examples display `n/a` for an absent count and preserve a reported zero. The accumulated equivalents are `modelMetrics[model].usage.cacheReadTokens` and `modelMetrics[model].usage.cacheWriteTokens` in `session.usage.getMetrics`.
+
+> [!NOTE]
+> These examples display the token categories separately. They do not define whether `inputTokens` includes cache reads or writes. Confirm that accounting convention with the runtime or provider before adding or subtracting these values, estimating charges, or comparing them with another SDK's token counts.
+
> [!TIP]
> `assistant.usage` is ephemeral, so it is delivered live but not replayed when you resume a session. To read accumulated totals after the fact, call `session.usage.getMetrics` (see [Accumulated AI credit and token totals](#accumulated-ai-credit-and-token-totals)).
@@ -68,8 +75,10 @@ const session = await client.createSession({ streaming: true });
session.on("assistant.usage", (event) => {
const { model, inputTokens, outputTokens, cost } = event.data;
+ const { cacheReadTokens, cacheWriteTokens } = event.data;
console.log(
- `${model}: in=${inputTokens ?? 0} out=${outputTokens ?? 0} cost=${cost ?? 0}`,
+ `${model}: in=${inputTokens ?? 0} out=${outputTokens ?? 0} ` +
+ `cache_read=${cacheReadTokens ?? "n/a"} cache_write=${cacheWriteTokens ?? "n/a"} cost=${cost ?? 0}`,
);
});
```
@@ -78,8 +87,10 @@ session.on("assistant.usage", (event) => {
```typescript
session.on("assistant.usage", (event) => {
const { model, inputTokens, outputTokens, cost } = event.data;
+ const { cacheReadTokens, cacheWriteTokens } = event.data;
console.log(
- `${model}: in=${inputTokens ?? 0} out=${outputTokens ?? 0} cost=${cost ?? 0}`,
+ `${model}: in=${inputTokens ?? 0} out=${outputTokens ?? 0} ` +
+ `cache_read=${cacheReadTokens ?? "n/a"} cache_write=${cacheWriteTokens ?? "n/a"} cost=${cost ?? 0}`,
);
});
```
@@ -100,7 +111,12 @@ session = await client.create_session(streaming=True)
def on_usage(event):
if event.type == SessionEventType.ASSISTANT_USAGE:
data = event.data
- print(f"{data.model}: in={data.input_tokens or 0} out={data.output_tokens or 0} cost={data.cost or 0}")
+ cache_read = data.cache_read_tokens if data.cache_read_tokens is not None else "n/a"
+ cache_write = data.cache_write_tokens if data.cache_write_tokens is not None else "n/a"
+ print(
+ f"{data.model}: in={data.input_tokens or 0} out={data.output_tokens or 0} "
+ f"cache_read={cache_read} cache_write={cache_write} cost={data.cost or 0}"
+ )
session.on(on_usage)
```
@@ -110,7 +126,12 @@ session.on(on_usage)
def on_usage(event):
if event.type == SessionEventType.ASSISTANT_USAGE:
data = event.data
- print(f"{data.model}: in={data.input_tokens or 0} out={data.output_tokens or 0} cost={data.cost or 0}")
+ cache_read = data.cache_read_tokens if data.cache_read_tokens is not None else "n/a"
+ cache_write = data.cache_write_tokens if data.cache_write_tokens is not None else "n/a"
+ print(
+ f"{data.model}: in={data.input_tokens or 0} out={data.output_tokens or 0} "
+ f"cache_read={cache_read} cache_write={cache_write} cost={data.cost or 0}"
+ )
session.on(on_usage)
```
@@ -159,7 +180,14 @@ func main() {
if d.Cost != nil {
cost = *d.Cost
}
- fmt.Printf("%s: in=%d out=%d cost=%g\n", d.Model, in, out, cost)
+ cacheRead, cacheWrite := "n/a", "n/a"
+ if d.CacheReadTokens != nil {
+ cacheRead = fmt.Sprint(*d.CacheReadTokens)
+ }
+ if d.CacheWriteTokens != nil {
+ cacheWrite = fmt.Sprint(*d.CacheWriteTokens)
+ }
+ fmt.Printf("%s: in=%d out=%d cache_read=%s cache_write=%s cost=%g\n", d.Model, in, out, cacheRead, cacheWrite, cost)
})
_ = session
}
@@ -182,7 +210,14 @@ session.On(func(event copilot.SessionEvent) {
if d.Cost != nil {
cost = *d.Cost
}
- fmt.Printf("%s: in=%d out=%d cost=%g\n", d.Model, in, out, cost)
+ cacheRead, cacheWrite := "n/a", "n/a"
+ if d.CacheReadTokens != nil {
+ cacheRead = fmt.Sprint(*d.CacheReadTokens)
+ }
+ if d.CacheWriteTokens != nil {
+ cacheWrite = fmt.Sprint(*d.CacheWriteTokens)
+ }
+ fmt.Printf("%s: in=%d out=%d cache_read=%s cache_write=%s cost=%g\n", d.Model, in, out, cacheRead, cacheWrite, cost)
})
```
@@ -202,7 +237,8 @@ session.On(evt =>
{
var data = evt.Data;
Console.WriteLine(
- $"{data.Model}: in={data.InputTokens ?? 0} out={data.OutputTokens ?? 0} cost={data.Cost ?? 0}");
+ $"{data.Model}: in={data.InputTokens ?? 0} out={data.OutputTokens ?? 0} " +
+ $"cache_read={data.CacheReadTokens?.ToString() ?? "n/a"} cache_write={data.CacheWriteTokens?.ToString() ?? "n/a"} cost={data.Cost ?? 0}");
});
```
@@ -212,7 +248,8 @@ session.On(evt =>
{
var data = evt.Data;
Console.WriteLine(
- $"{data.Model}: in={data.InputTokens ?? 0} out={data.OutputTokens ?? 0} cost={data.Cost ?? 0}");
+ $"{data.Model}: in={data.InputTokens ?? 0} out={data.OutputTokens ?? 0} " +
+ $"cache_read={data.CacheReadTokens?.ToString() ?? "n/a"} cache_write={data.CacheWriteTokens?.ToString() ?? "n/a"} cost={data.Cost ?? 0}");
});
```
@@ -228,7 +265,9 @@ session.on(AssistantUsageEvent.class, event -> {
long in = data.inputTokens() != null ? data.inputTokens() : 0;
long out = data.outputTokens() != null ? data.outputTokens() : 0;
double cost = data.cost() != null ? data.cost() : 0.0;
- System.out.printf("%s: in=%d out=%d cost=%s%n", data.model(), in, out, cost);
+ String cacheRead = data.cacheReadTokens() != null ? data.cacheReadTokens().toString() : "n/a";
+ String cacheWrite = data.cacheWriteTokens() != null ? data.cacheWriteTokens().toString() : "n/a";
+ System.out.printf("%s: in=%d out=%d cache_read=%s cache_write=%s cost=%s%n", data.model(), in, out, cacheRead, cacheWrite, cost);
});
```
@@ -244,11 +283,17 @@ let mut events = session.subscribe();
while let Ok(event) = events.recv().await {
if event.event_type == "assistant.usage" {
if let Some(data) = event.typed_data::() {
+ let cache_read = data.cache_read_tokens
+ .map(|tokens| tokens.to_string()).unwrap_or_else(|| "n/a".to_string());
+ let cache_write = data.cache_write_tokens
+ .map(|tokens| tokens.to_string()).unwrap_or_else(|| "n/a".to_string());
println!(
- "{}: in={} out={} cost={}",
+ "{}: in={} out={} cache_read={} cache_write={} cost={}",
data.model,
data.input_tokens.unwrap_or(0),
data.output_tokens.unwrap_or(0),
+ cache_read,
+ cache_write,
data.cost.unwrap_or(0.0),
);
}
@@ -672,7 +717,7 @@ The example uses the fields below. The generated `UsageGetMetricsResult` type is
|---|---|---|
| `totalNanoAiu` | `number` | Session-wide AI credit cost, in nano-AI units |
| `totalPremiumRequestCost` | `number` | Premium request cost across all models, after multipliers |
-| `modelMetrics` | `Record` | Per-model breakdown; each entry has `usage.inputTokens`, `usage.outputTokens`, and `totalNanoAiu` |
+| `modelMetrics` | `Record` | Per-model breakdown; each entry has `usage.inputTokens`, `usage.outputTokens`, `usage.cacheReadTokens`, `usage.cacheWriteTokens`, and `totalNanoAiu` |
> [!NOTE]
> Cost is reported in **nano-AI units** (the field is named `totalNanoAiu`). The exact conversion to AI credits and the precise meaning of premium request accounting are defined by GitHub Copilot billing, not by the SDK—treat [GitHub's Copilot billing documentation](https://docs.github.com/en/copilot/managing-copilot/understanding-and-managing-copilot-usage) as the source of truth and verify before surfacing currency-like values to users. The examples divide by `1e9` as a convenience, following the SI `nano` prefix; confirm this matches current billing before relying on it. The `modelMetrics` and `tokenDetails` maps are keyed by runtime strings (model IDs and token-type names) that the SDK type system does not validate.