Repository navigation
Conversation
Implement the handler routing rules of TESTING.md §8 in launchdarkly/ai-sdks-monorepo#52. - run() takes `handler` or `handlers`, not both. `judge_handlers` is removed. Every handler must declare provides_for, and two handlers cannot declare the same provider and mode. - Generation and every judge select a handler with the config() rule (utils.select_handler). There is no mode fallback, so a messages-mode judge never runs on an agent handler. - generation.mode is required. The values are "completion" and "agent". GenerationConfig stays a TypedDict with Required fields. GenerationOverrides is the type for overrides with ai_config. - The prompt must match the handler mode. A messages handler needs messages and an agent handler needs instructions. The harness does not convert one into the other. - Judges get no tools. - ai_config reads the AI Config mode first, then the variation, then the model config. Each check runs right after the read that supplies its input. The mode selects the prompt field of the variation. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
| selected = select_handler( | ||
| {"provider": {"name": provider}}, | ||
| {"mode": mode}, | ||
| list(handlers), # type: ignore[arg-type] | ||
| ) |
There was a problem hiding this comment.
🟡 List-tagged handlers never match their provider
When provides_for is a list, _validate_handlers accepts it, but select_handler compares the original list against a tuple. Valid custom handlers fail routing, so evaluation runs stop before generation.
Learn more
Handler metadata can be a two-element list or tuple, and _validate_handlers accepts either through _provides_for. The shared select_handler compares the original metadata directly to tuples. Consequently a list-tagged handler passes validation but is never selected for its own provider and mode. An exact matching handler is necessary before a generation run can proceed.
Example: With a sole handler declaring provides_for = ["OpenAI", "messages"] and generation set to {"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"}, validation succeeds but routing raises No handler can run generation.
Recommended fix: Normalize accepted handler metadata to tuples before passing candidates to select_handler, or reject list metadata consistently at validation if the public contract is changed. Test generation and judge selection using list-tagged custom handlers.
Was this helpful? React with 👍 or 👎 to provide feedback.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 26eafe0. Configure here.
| ) | ||
| except ValueError: | ||
| return None | ||
| return cast(EvalHandler, selected) |
There was a problem hiding this comment.
Handler mode aliases fail selection
Medium Severity
select_eval_handler matches handlers through select_handler, which compares raw provides_for tuples, while _provides_for and _validate_handlers treat completion and messages as the same mode. A handler declared as (provider, "completion") is accepted as covering messages mode, then selection cannot find it. The error from describe_handlers then lists the normalized messages pair, so the pool appears to already have the required handler.
Additional Locations (2)
Reviewed by Cursor Bugbot for commit 26eafe0. Configure here.


Summary
This change implements the handler routing rules of the shared spec,
TESTING.md§8, in launchdarkly/ai-sdks-monorepo#52. It lets one evaluation run use different providers or modes for generation and for judges. For example, generation uses OpenAI and a judge uses Anthropic, or generation uses an agent handler and a judge uses a messages handler.This is a breaking change to the experimental
evaluationsAPI.handlerorhandlersrun()takes one handler ashandleror a list ashandlers, not both.judge_handlersis removed, with no alias.provides_for. Before this change, a handler without it received every judge config with no warning.One selection rule
utils.select_handler, the same rule asconfig(). The evals code has no copy of it.collapse_messagesis removed from the evals path. The onlinejudges.pypath does not change.generation.modemodeis required:"completion"or"agent". There is no default."messages","judge"or a typo, raise before any network I/O. They do not go throughnormalize_mode.GenerationConfigstays aTypedDictwithRequired[...], so callers pass a plaindictand import no type. The newGenerationOverridesTypedDictis the type for overrides withai_config. Two@overloadsignatures onrun()select between them.GenerationConfigdocstring and the README tell the user how to choose a mode.Prompt must match the handler mode
instructionsraises: "the generation prompt needsmessages".messagesraises: "the generation prompt needsinstructions".Judges get no tools
A judge call gets an empty tool map, and
toolsis removed from the judge config.ai_configGET ai-configs/<key>to get the mode. The mode is on the AI Config, not on the variation.instructionsin agent mode,messagesin completion and judge mode.mode.Migration
To evaluate your own code, wrap it:
create_handler(("OpenAI", "messages"), my_app).Tests
mode. The judge fixtures now usemessages, which is what judge configs carry.ai_configfailuremake test: 1565 passed, 11 skipped.make lint,make format-checkandmake typecheckare clean.Out of scope
toolsmap shape and itsGET ai-toolsstep at run time. Python useslist[EvalTool]andtools.get(), which was an earlier divergence.🤖 Generated with Claude Code
Note
Overview
Breaking change to the experimental offline evaluations API: handler routing now follows the same provider-and-mode rules as
config(), with stricter validation up front.run()acceptshandlerorhandlers(mutually exclusive);judge_handlersis removed. Every handler must declareprovides_for, and duplicate provider/mode pairs are rejected. Generation and each judge pick a handler viaselect_handler—no fallback (e.g. a messages-mode judge never runs on an agent handler; message collapsing for judges is gone).generation.modeis required ("completion"or"agent", no default). The prompt must match the selected handler:messagesfor completion/messages handlers,instructionsfor agent handlers; mismatches fail before any network I/O. Judge resolution aggregates all handler/prompt problems into one error before records are created. Judges receive no tools (empty tool map,toolsstripped from judge config).ai_configruns now read the AI Config mode first (GET ai-configs/<key>), then variation and model config, validating handler coverage, prompt, and provider after each step.GenerationOverridesallows field overrides but notmode. New exports:GenerationMode,GenerationOverrides. READMEs and tests are updated formode, tagged handlers, and the unifiedhandlerslist.Reviewed by Cursor Bugbot for commit 26eafe0. Bugbot is set up for automated code reviews on this repo. Configure here.