Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions packages/ai/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,16 +63,16 @@ evals = init_evaluations(project_key="my-project")
result = await evals.run(
key="unique-evaluation-key",
dataset="golden-dataset",
handler=my_handler,
generation={"provider": "OpenAI", "model": "gpt-4o"},
handler=my_handler, # built with create_handler(("OpenAI", "messages"), fn)
generation={"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"},
criteria=[
Judge(key="accuracy-judge"),
Scorer(name="mentions-policy", fn=lambda row, output: "policy" in (output or "")),
],
)
```

`LD_API_TOKEN` is required. Configure `LD_SDK_KEY` — or initialize your own client with `init_client(client=...)` — to emit one `$ld:ai:offline-evals:generation` event per generated row, plus one `$ld:ai:offline-evals:criterion` event per `(row, criterion)` when `criteria` are supplied, through the standard SDK event transport. The SDK reports scores; LaunchDarkly rules on them at ingest. A judge served by a different provider than `generation` needs a handler for it in `judge_handlers`. Each row's tool calls are recorded during generation and rendered into the judge's `{{message_history}}`, between the row input and the generated output, so a rubric can grade the tool trajectory as well as the final answer. Use `LD_API_BASE_URI` for staging or local management API traffic; it is separate from the SDK delivery setting `LD_BASE_URI`. Evaluation-run links use the explicit `ui_base_uri` option or `LD_UI_BASE_URI`, defaulting to `https://app.launchdarkly.com`; set it when the project is not in production, or a run created elsewhere still links to the production app. `tools` is a list of `EvalTool`. Construct one to define a tool in code, or await `evals.tools.get()` for a tool that already exists in LaunchDarkly. See the [core evaluations guide](https://github.com/launchdarkly/python-ai-sdk/blob/main/packages/client/README.md#run-an-evaluation-from-code).
`LD_API_TOKEN` is required. Configure `LD_SDK_KEY` — or initialize your own client with `init_client(client=...)` — to emit one `$ld:ai:offline-evals:generation` event per generated row, plus one `$ld:ai:offline-evals:criterion` event per `(row, criterion)` when `criteria` are supplied, through the standard SDK event transport. The SDK reports scores; LaunchDarkly rules on them at ingest. Pass one handler as `handler` or a list as `handlers`. The generation config and each judge select a handler by provider and mode, so a judge served by a different provider than `generation` needs its own handler in `handlers`. `generation.mode` is required: `"completion"` for a messages handler and a `messages` prompt, `"agent"` for an agent handler and an `instructions` prompt. Each row's tool calls are recorded during generation and rendered into the judge's `{{message_history}}`, between the row input and the generated output, so a rubric can grade the tool trajectory as well as the final answer. Use `LD_API_BASE_URI` for staging or local management API traffic; it is separate from the SDK delivery setting `LD_BASE_URI`. Evaluation-run links use the explicit `ui_base_uri` option or `LD_UI_BASE_URI`, defaulting to `https://app.launchdarkly.com`; set it when the project is not in production, or a run created elsewhere still links to the production app. `tools` is a list of `EvalTool`. Construct one to define a tool in code, or await `evals.tools.get()` for a tool that already exists in LaunchDarkly. See the [core evaluations guide](https://github.com/launchdarkly/python-ai-sdk/blob/main/packages/client/README.md#run-an-evaluation-from-code).

---

Expand Down
30 changes: 19 additions & 11 deletions packages/client/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,7 +67,8 @@ async def main() -> int:
generation={
"provider": "OpenAI",
"model": "gpt-4o",
"instructions": "You are a support agent.",
"mode": "completion",
"messages": [{"role": "system", "content": "You are a support agent."}],
},
)
print(result.url, result.summary)
Expand All @@ -77,7 +78,15 @@ async def main() -> int:
sys.exit(asyncio.run(main()))
```

`project_key` is supplied during initialization, either as an argument to `init_evaluations()` or through `LD_PROJECT_KEY`. `generation.instructions` is shorthand for one system message; use `generation.messages` instead for a full message list, but do not supply both. The harness never retries a handler invocation because doing so could repeat tool side effects. Its retries apply only to LaunchDarkly management API requests.
`project_key` is supplied during initialization, either as an argument to `init_evaluations()` or through `LD_PROJECT_KEY`. `generation` requires `provider`, `model`, and `mode`. `generation` is a `TypedDict`, so you pass a plain `dict` and import no type.

**Choose the mode.** Use `"completion"` when the prompt is a list of `messages` and the handler sends one request to a chat or messages API, for example `create_openai_messages_handler()`. Use `"agent"` when the prompt is one `instructions` string and an agent framework runs the model, for example OpenAI Agents, Claude Agents, or LangChain agents. If you are not sure, use the mode of the AI Config in LaunchDarkly. The mode has no default.

**The prompt must match the handler.** A messages handler needs `messages`, and an agent handler needs `instructions`. The harness does not convert one into the other: a mismatch raises before any network I/O, and the error names the field the handler needs.

**Pass one handler as `handler` or a list as `handlers`, not both.** Each handler must declare `provides_for`; the provider packages' `create_*_handler()` factories and `create_handler()` set it. The generation config and every judge select a handler from the list by provider and mode, with the same rule as `config()`: a handler that names the provider wins over a `"*"` handler, and the mode must match exactly. To evaluate your own application code, wrap it with `create_handler((provider, mode), fn)`.

**Start from an AI Config.** Pass `ai_config=AIConfig(key=..., variation=...)` instead of a full `generation`. The harness reads the AI Config's mode first, then the variation, then the model config it links, and runs each check as soon as the read that supplies its input returns. The mode selects the prompt field: `instructions` in agent mode, `messages` in completion and judge mode. `generation` can then override single fields, except `mode`. The harness never retries a handler invocation because doing so could repeat tool side effects. Its retries apply only to LaunchDarkly management API requests.

Generation and criterion events are the only path by which row results reach LaunchDarkly, so `init_evaluations()` raises rather than creating a run that can never complete unless it can resolve an event transport: either an SDK key (`sdk_key` or `LD_SDK_KEY`) or a client already initialized through `init_client(client=...)`. Bringing your own client lets a process emit evaluation events without an SDK key in scope. Every generated row is emitted and flushed unconditionally; no feature flag gates event publishing. The harness then polls the summary endpoint until row accounting shows processing is complete.

Expand All @@ -94,7 +103,7 @@ result = await evals.run(
{"input": "Where is order {{order_id}}?", "variables": {"order_id": "A-17"}},
],
handler=create_openai_messages_handler(),
generation={"provider": "OpenAI", "model": "gpt-4o"},
generation={"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"},
)
```

Expand All @@ -117,15 +126,14 @@ def mentions_policy(row: DatasetRow, output: str | None) -> bool:
result = await init_evaluations(project_key="my-project").run(
key="support-qa-2026-08-20",
dataset="support-golden",
handler=create_openai_messages_handler(),
generation={"provider": "OpenAI", "model": "gpt-4o"},
# The generation config uses OpenAI and the judge uses Anthropic, so the
# list holds a handler for each.
handlers=[create_openai_messages_handler(), create_claude_messages_handler()],
generation={"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"},
criteria=[
Judge(key="accuracy-judge", threshold=0.8),
Scorer(name="mentions-policy", fn=mentions_policy),
],
# Needed only because this judge is served by a different provider than
# the generation config above.
judge_handlers=[create_claude_messages_handler()],
)
```

Expand Down Expand Up @@ -178,9 +186,9 @@ Two limits keep a trajectory from spending the judge's context window: at most 5

A tool result is now judge-prompt input. It stays literal for the same reason the generated output does: the judge config is handed to the handler unrendered and the handler makes exactly one template pass, so a `{{...}}` sequence coming back from a tool is never expanded into the judge prompt.

**Judges are independent AI Configs, so handlers are routed per judge.** A judge may resolve to a different provider or mode than `generation`, and a handler built for one provider cannot execute another's config. `handler` runs a judge when it provides for that judge's provider; pass handlers for any other providers in `judge_handlers`. Selection prefers a handler naming the judge's provider outright over a wildcard multi-provider adapter, and an agent-mode handler can serve a messages-mode judge with its messages collapsed into one instructions block. A plain callable that declares no `provides_for` routes itself, exactly as it already does for the generation config.
**Judges are independent AI Configs, so handlers are routed per judge.** A judge may resolve to a different provider or mode than `generation`. Each judge selects its handler from `handler` or `handlers` by its provider and mode, with the same rule as generation. There is no mode fallback: a messages-mode judge never runs on an agent handler. A judge's prompt must also match its handler's mode. Judges get no tools: the harness passes an empty tool map and removes `tools` from the judge config.

Judges are resolved through flag delivery, and handlers are matched to them, **before** any evaluation records are created — a missing judge or one no handler covers fails the run up front rather than after the generation spend. After that point a criterion failure never aborts the run: an unparseable judge response, an out-of-range score, a raising handler or scorer, and a row whose generation errored each become a per-criterion `ERROR` event with a cause code (`invalid_judge_output`, `invalid_score`, `handler_raised`, `scorer_raised`, `generation_incomplete`) and a top-level `errorMessage`. Event *delivery* is different: the backend needs one result per `(row, criterion)` to finish row accounting, so if tracking a criterion event fails, every remaining result is still attempted and flushed and then `run()` raises — rather than polling to its timeout with the cause hidden.
Judges are resolved through flag delivery, and handlers are matched to them, **before** any evaluation records are created — a missing judge, a judge that no handler covers, or a judge whose prompt does not match its handler fails the run up front rather than after the generation spend. One error lists every judge with a problem. After that point a criterion failure never aborts the run: an unparseable judge response, an out-of-range score, a raising handler or scorer, and a row whose generation errored each become a per-criterion `ERROR` event with a cause code (`invalid_judge_output`, `invalid_score`, `handler_raised`, `scorer_raised`, `generation_incomplete`) and a top-level `errorMessage`. Event *delivery* is different: the backend needs one result per `(row, criterion)` to finish row accounting, so if tracking a criterion event fails, every remaining result is still attempted and flushed and then `run()` raises — rather than polling to its timeout with the cause hidden.

The client uses **lazy initialization**: importing the package does not connect to LaunchDarkly. The singleton is created automatically on the first API call that needs it (`config().invoke()`, `graph().invoke()`, `resolve_graph()`, etc.), as long as `LD_SDK_KEY` is set in the environment.

Expand Down Expand Up @@ -250,7 +258,7 @@ result = await evals.run(
key="support-qa-2026-08-20",
dataset="support-golden",
handler=create_openai_messages_handler(),
generation={"provider": "OpenAI", "model": "gpt-4o"},
generation={"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"},
tools=[search_docs_tool, lookup_order_tool],
)
```
Expand Down
Loading
Loading