Skip to content

feat(evaluations)!: route handlers by provider and mode - #152

Open
donei003 wants to merge 1 commit into
mainfrom
claude/evals-handler-routing
Open

donei003 wants to merge 1 commit into
mainfrom
claude/evals-handler-routing

Conversation

@donei003

@donei003 donei003 commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Summary

This change implements the handler routing rules of the shared spec, TESTING.md §8, in launchdarkly/ai-sdks-monorepo#52. It lets one evaluation run use different providers or modes for generation and for judges. For example, generation uses OpenAI and a judge uses Anthropic, or generation uses an agent handler and a judge uses a messages handler.

This is a breaking change to the experimental evaluations API.

handler or handlers

  • run() takes one handler as handler or a list as handlers, not both. judge_handlers is removed, with no alias.
  • Every handler must declare provides_for. Before this change, a handler without it received every judge config with no warning.
  • Two handlers with the same provider and mode raise an error.

One selection rule

  • Generation and every judge select a handler with utils.select_handler, the same rule as config(). The evals code has no copy of it.
  • There is no mode fallback. A messages-mode judge never runs on an agent handler, so collapse_messages is removed from the evals path. The online judges.py path does not change.

generation.mode

  • mode is required: "completion" or "agent". There is no default.
  • Other values, such as "messages", "judge" or a typo, raise before any network I/O. They do not go through normalize_mode.
  • GenerationConfig stays a TypedDict with Required[...], so callers pass a plain dict and import no type. The new GenerationOverrides TypedDict is the type for overrides with ai_config. Two @overload signatures on run() select between them.
  • The GenerationConfig docstring and the README tell the user how to choose a mode.
  • The mode is not sent to LaunchDarkly.

Prompt must match the handler mode

  • A messages handler with instructions raises: "the generation prompt needs messages".
  • An agent handler with messages raises: "the generation prompt needs instructions".
  • The same check applies to each judge. All judge problems are reported in one error, before any record is created.

Judges get no tools

A judge call gets an empty tool map, and tools is removed from the judge config.

ai_config

  • The run first reads GET ai-configs/<key> to get the mode. The mode is on the AI Config, not on the variation.
  • Then it reads the variation, and then the model config.
  • Each check runs right after the read that supplies its input:
    • mode coverage: after 1 GET
    • prompt: after 2 GETs
    • provider: after the model config read
  • The mode selects the variation's prompt field: instructions in agent mode, messages in completion and judge mode.
  • When the mode's field is empty, the other field is kept, so that the prompt check rejects it.
  • Overrides cannot set mode.

Migration

# Before
await evals.run(..., handler=openai_handler, judge_handlers=[claude_handler],
                generation={"provider": "OpenAI", "model": "gpt-4o"})

# After
await evals.run(..., handlers=[openai_handler, claude_handler],
                generation={"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"})

To evaluate your own code, wrap it: create_handler(("OpenAI", "messages"), my_app).

Tests

  • The existing tests now use tagged handlers and set mode. The judge fixtures now use messages, which is what judge configs carry.
  • New tests cover:
    • handler argument validation
    • mode validation
    • prompt and mode mismatches
    • selection by provider and mode, and exact provider over wildcard
    • no mode fallback for judges
    • one error for all judge problems
    • judges with no tools
    • the request count at each ai_config failure
    • prompt field selection by mode
  • make test: 1565 passed, 11 skipped. make lint, make format-check and make typecheck are clean.

Out of scope

  • The spec's tools map shape and its GET ai-tools step at run time. Python uses list[EvalTool] and tools.get(), which was an earlier divergence.

🤖 Generated with Claude Code


Note

Overview
Breaking change to the experimental offline evaluations API: handler routing now follows the same provider-and-mode rules as config(), with stricter validation up front.

run() accepts handler or handlers (mutually exclusive); judge_handlers is removed. Every handler must declare provides_for, and duplicate provider/mode pairs are rejected. Generation and each judge pick a handler via select_handler—no fallback (e.g. a messages-mode judge never runs on an agent handler; message collapsing for judges is gone).

generation.mode is required ("completion" or "agent", no default). The prompt must match the selected handler: messages for completion/messages handlers, instructions for agent handlers; mismatches fail before any network I/O. Judge resolution aggregates all handler/prompt problems into one error before records are created. Judges receive no tools (empty tool map, tools stripped from judge config).

ai_config runs now read the AI Config mode first (GET ai-configs/<key>), then variation and model config, validating handler coverage, prompt, and provider after each step. GenerationOverrides allows field overrides but not mode. New exports: GenerationMode, GenerationOverrides. READMEs and tests are updated for mode, tagged handlers, and the unified handlers list.

Reviewed by Cursor Bugbot for commit 26eafe0. Bugbot is set up for automated code reviews on this repo. Configure here.

Implement the handler routing rules of TESTING.md §8 in
launchdarkly/ai-sdks-monorepo#52.

- run() takes `handler` or `handlers`, not both. `judge_handlers` is
  removed. Every handler must declare provides_for, and two handlers
  cannot declare the same provider and mode.
- Generation and every judge select a handler with the config() rule
  (utils.select_handler). There is no mode fallback, so a messages-mode
  judge never runs on an agent handler.
- generation.mode is required. The values are "completion" and "agent".
  GenerationConfig stays a TypedDict with Required fields.
  GenerationOverrides is the type for overrides with ai_config.
- The prompt must match the handler mode. A messages handler needs
  messages and an agent handler needs instructions. The harness does not
  convert one into the other.
- Judges get no tools.
- ai_config reads the AI Config mode first, then the variation, then the
  model config. Each check runs right after the read that supplies its
  input. The mode selects the prompt field of the variation.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

Devin Review

Comment on lines +117 to +121
selected = select_handler(
{"provider": {"name": provider}},
{"mode": mode},
list(handlers), # type: ignore[arg-type]
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 List-tagged handlers never match their provider

When provides_for is a list, _validate_handlers accepts it, but select_handler compares the original list against a tuple. Valid custom handlers fail routing, so evaluation runs stop before generation.

Learn more

Handler metadata can be a two-element list or tuple, and _validate_handlers accepts either through _provides_for. The shared select_handler compares the original metadata directly to tuples. Consequently a list-tagged handler passes validation but is never selected for its own provider and mode. An exact matching handler is necessary before a generation run can proceed.

Example: With a sole handler declaring provides_for = ["OpenAI", "messages"] and generation set to {"provider": "OpenAI", "model": "gpt-4o", "mode": "completion"}, validation succeeds but routing raises No handler can run generation.

Recommended fix: Normalize accepted handler metadata to tuples before passing candidates to select_handler, or reject list metadata consistently at validation if the public contract is changed. Test generation and judge selection using list-tagged custom handlers.

Devin Review


Was this helpful? React with 👍 or 👎 to provide feedback.

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, have a team admin enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 26eafe0. Configure here.

)
except ValueError:
return None
return cast(EvalHandler, selected)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Handler mode aliases fail selection

Medium Severity

select_eval_handler matches handlers through select_handler, which compares raw provides_for tuples, while _provides_for and _validate_handlers treat completion and messages as the same mode. A handler declared as (provider, "completion") is accepted as covering messages mode, then selection cannot find it. The error from describe_handlers then lists the normalized messages pair, so the pool appears to already have the required handler.

Additional Locations (2)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 26eafe0. Configure here.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant