Skip to content

Add advanced BCal evaluation categories - #890

Draft
Esben Nyhuus Kristoffersen (esbenk) wants to merge 3 commits into
mainfrom
esbenk-bcal-advanced-evals
Draft

Esben Nyhuus Kristoffersen (esbenk) wants to merge 3 commits into
mainfrom
esbenk-bcal-advanced-evals

Conversation

@esbenk

Copy link
Copy Markdown
Collaborator

BC-Bench needs scenario-native evaluation for BCal's multi-turn interaction, planning, and feature-development behavior instead of extending the single-prompt nl2al category. This adds two categories that share a strongly validated runner and evaluation layer while keeping the provisional BCal integration isolated.

What changed

  • Adds bcal-scenario and bcal-feature dataset schemas plus three pilot entries covering integration choice, explicit plan-first lifecycle, and warehouse inventory-risk management.
  • Generates redacted BCal execution manifests, invokes the finalized scenario CLI contract, parses structured results and sanitized session archives, captures token/turn/tool metrics, and exports the generated workspace as an artifact.
  • Adds deterministic trace evaluation for tool requirements/order, interaction matching, plan lifecycle, compile-after-mutation, and forbidden publishing behavior.
  • Adds structured scenario results and bc-eval metrics for scenario completion, independent build, and trace compliance while retaining LMChecklist as the core score.
  • Sends LMChecklist a rich packet containing the task, confirmed decisions, deterministic and trace summaries, workspace diff, and complete textual export.
  • Wires category mappings, CLI restrictions, PowerShell dataset lookup, mock evaluation, BCal workflow inputs, documentation, and exhaustiveness/integrity tests.

Validation

  • ruff check .
  • Targeted ty check for changed Python modules
  • Full test suite: 1064 passed, 2 skipped, 1 deselected
  • BCal workflow YAML parsing and PowerShell category-to-dataset mappings

Review note

Independent compilation uses standalone al compile when al is available on PATH. The current symbols-only workflow explicitly records not_attempted with a concrete reason and does not count it as a successful independent build. Runtime verification remains N/A for the initial pilots.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
import pytest
from pydantic import ValidationError

import bcbench.evaluate.bcal_scenario as scenario_evaluation

from bcbench.agent.bcal.scenario import SCENARIO_EXPORT_DIR, SCENARIO_RESULT, SCENARIO_SYMBOL_DIR, BCalScenarioRunResult, load_scenario_result, resolve_session_chat
from bcbench.dataset import BCalScenarioEntry, BCalTraceAssertion
from bcbench.evaluate.base import AgentRunner, EvaluationPipeline
martinsrui-msft and others added 2 commits September 25, 2026 13:52
The bcal-scenario and bcal-feature runs failed in ~3s, before any LLM call,
because the session block sent values that bcal's scenario schema rejects:

  The JSON value could not be converted to
  System.Nullable`1[Microsoft.BusinessCentral.BCal.Service.BcalAppMode].
  Path: $.session.mode

Two fields were wrong:

- "mode" was "agent"/"plan". BcalAppMode is extension | customization |
  personalization -- it selects the kind of app bcal builds, not the agent's
  interaction style. All three entries build AL extensions, so they now use
  "extension" (also bcal's own default). Plan-first behavior is already
  covered by the plan_action step and the plan_lifecycle trace assertions,
  which is the correct mechanism.
- "publish" was the boolean false, but BcalPublishMode is never | ask |
  always. This would have failed immediately after the mode fix. The intent
  (all three assert the no-publish forbidden_tool) maps to "never".

Both fields are now constrained with Literal in BCalSessionConfiguration, so
an invalid value fails dataset validation with the accepted values listed,
instead of surfacing as an opaque System.Text.Json error mid-run.

Verified: full suite 1064 passed, and both datasets load via
`bcbench dataset list`.

Note: session.resume is bool | None here but string? on the bcal side. It is
excluded when unset so it cannot break today, left unchanged as out of scope.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0875f4b-d1d0-4ded-893f-db73c0c3f8f5
CI's lint-and-test job runs `pre-commit run --all-files`, whose ruff-format
hook rewrites files and then fails because the tree was dirty. Two files were
left unformatted:

- tests/test_bcal_scenario.py: ruff normalizes the outer quotes on the
  --category assertion to double quotes (same number of escapes, so double
  wins).
- tests/test_expected_metrics_warning.py: the expected-warning list fits on a
  single line under the configured line length, so the manual wrap is undone.

Both violations pre-date this branch's dataset fix -- lint-and-test was already
failing on 24062b0 for the same reason. Formatting only; no behavior change.

Verified: `pre-commit run --all-files` passes all 11 hooks, and the suite is
1064 passed / 2 skipped.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: e0875f4b-d1d0-4ded-893f-db73c0c3f8f5

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants