Add advanced BCal evaluation categories - #890
Draft
Esben Nyhuus Kristoffersen (esbenk) wants to merge 3 commits into
Draft
Esben Nyhuus Kristoffersen (esbenk) wants to merge 3 commits into
Esben Nyhuus Kristoffersen (esbenk) wants to merge 3 commits into
Conversation
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
| import pytest | ||
| from pydantic import ValidationError | ||
|
|
||
| import bcbench.evaluate.bcal_scenario as scenario_evaluation |
|
|
||
| from bcbench.agent.bcal.scenario import SCENARIO_EXPORT_DIR, SCENARIO_RESULT, SCENARIO_SYMBOL_DIR, BCalScenarioRunResult, load_scenario_result, resolve_session_chat | ||
| from bcbench.dataset import BCalScenarioEntry, BCalTraceAssertion | ||
| from bcbench.evaluate.base import AgentRunner, EvaluationPipeline |
The bcal-scenario and bcal-feature runs failed in ~3s, before any LLM call, because the session block sent values that bcal's scenario schema rejects: The JSON value could not be converted to System.Nullable`1[Microsoft.BusinessCentral.BCal.Service.BcalAppMode]. Path: $.session.mode Two fields were wrong: - "mode" was "agent"/"plan". BcalAppMode is extension | customization | personalization -- it selects the kind of app bcal builds, not the agent's interaction style. All three entries build AL extensions, so they now use "extension" (also bcal's own default). Plan-first behavior is already covered by the plan_action step and the plan_lifecycle trace assertions, which is the correct mechanism. - "publish" was the boolean false, but BcalPublishMode is never | ask | always. This would have failed immediately after the mode fix. The intent (all three assert the no-publish forbidden_tool) maps to "never". Both fields are now constrained with Literal in BCalSessionConfiguration, so an invalid value fails dataset validation with the accepted values listed, instead of surfacing as an opaque System.Text.Json error mid-run. Verified: full suite 1064 passed, and both datasets load via `bcbench dataset list`. Note: session.resume is bool | None here but string? on the bcal side. It is excluded when unset so it cannot break today, left unchanged as out of scope. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0875f4b-d1d0-4ded-893f-db73c0c3f8f5
CI's lint-and-test job runs `pre-commit run --all-files`, whose ruff-format hook rewrites files and then fails because the tree was dirty. Two files were left unformatted: - tests/test_bcal_scenario.py: ruff normalizes the outer quotes on the --category assertion to double quotes (same number of escapes, so double wins). - tests/test_expected_metrics_warning.py: the expected-warning list fits on a single line under the configured line length, so the manual wrap is undone. Both violations pre-date this branch's dataset fix -- lint-and-test was already failing on 24062b0 for the same reason. Formatting only; no behavior change. Verified: `pre-commit run --all-files` passes all 11 hooks, and the suite is 1064 passed / 2 skipped. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: e0875f4b-d1d0-4ded-893f-db73c0c3f8f5
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
BC-Bench needs scenario-native evaluation for BCal's multi-turn interaction, planning, and feature-development behavior instead of extending the single-prompt
nl2alcategory. This adds two categories that share a strongly validated runner and evaluation layer while keeping the provisional BCal integration isolated.What changed
bcal-scenarioandbcal-featuredataset schemas plus three pilot entries covering integration choice, explicit plan-first lifecycle, and warehouse inventory-risk management.Validation
ruff check .ty checkfor changed Python modules1064 passed, 2 skipped, 1 deselectedReview note
Independent compilation uses standalone
al compilewhenalis available on PATH. The current symbols-only workflow explicitly recordsnot_attemptedwith a concrete reason and does not count it as a successful independent build. Runtime verification remains N/A for the initial pilots.