knowledge: improve review precision from BCApps PR 10990-neg-8222ad38e67a1bf4ee3b67d2389a957f0d00c77a feedback - #206
Conversation
…e67a1bf4ee3b67d2389a957f0d00c77a feedback
Jesper Schulz-Wedde (JesperSchulz)
left a comment
There was a problem hiding this comment.
Reviewed draft head 3966f942976d4cadf7304047635f9ca4997c09a7. The exception is currently justified by a false regeneration model: the article says Contoso/demo modules are explicitly rerun and data is regenerated on demand, but Contoso Demo Tool records generated modules via Data Level, excludes modules for which IsModuleGenerated(...) is true, and errors when selected modules were already generated. Generation also occurs through new-company/build entry points, so it is not exclusively manual. This could suppress legitimate upgrade findings based on a nonexistent rerun path.
Please scope the exception to data explicitly owned by a disposable/unsupported demo-data module where compatibility for previously generated demo companies is intentionally not required. Remove claims that modules are rerun or regenerated on demand, and do not treat generic preview/trial status as proof that persisted data needs no upgrade path. Add a negative evaluation fixture if this suppression must remain deterministic.
Summary
Improves BCQuality knowledge based on maintainer thumbs-down feedback from explicit BCApps PR review runs.
Source feedback
Validation
Generated by the BC-ALAgentsInternal self-improvement workflow.
Offline evaluation: regression
Candidate correctness failed: unexpected or missed findings remain. Existing ignored gold comments retain their neutral scoring semantics.
Preparation attempt 1: validated.
Selection: new. New:
synthetic__upgrade-demo-data-not-upgraded-01. Reused (complete payloads): ``.synthetic__upgrade-demo-data-not-upgraded-01/calibration_context/ Add-Expense-VAT-settings-to-Contoso-demo-tool BCApps#10990 (comment): Reproduces the reviewed boundary: a Contoso demo-data module's CreateMasterData() wires a new VAT-rate seeding codeunit that inserts persisted setup records, with no Subtype = Upgrade codeunit to propagate those records to already-provisioned tenants, matching the bot's upgrade-codeunit-subtype and install-code-does-not-run-on-version-upgrade reasoning. Feedback was negative (THUMBS_DOWN) and the author replied 'Not needed here, this is Preview demo data', rejecting applicability to Contoso/demo-data tooling specifically -- a product/process judgment call, not a confirmed absence of the underlying defect pattern. Per the no-invented-clean-negative rule, the expected finding is retained rather than converted into a false_positive_guard; this is deliberately scoped as calibration_context so the disagreement is captured honestly. Limitation: does not prove the bot's generic upgrade-safety guidance extends to demo-data/Preview-tooling modules, and does not resolve whether BC's Contoso demo-data lifecycle exempts such modules from Subtype = Upgrade requirements -- that remains an open, human-reviewed question.Coverage references and outcome shapes are validated mechanically. Semantic equivalence, severity calibration and recommendation quality are NOT proved by these checks or a matching F1.
Common engine code:
ecf8e31759d6ddd6d78e3a0b7836b40134368009; dataset SHA256:46A978950BE2F3B98ED647D3AF503448E74879D29E95988C1394EB90DD443EB6.Model:
gpt-5.6-luna; judge:gpt-5.3-codex. Exact D:synthetic__upgrade-demo-data-not-upgraded-01.Dataset base:
e89b01069c079e3e4741587ffd5b4c45eae85143; candidate:ddb215cb28634e8332a2131f8c992833b718a150.d537505e3e236358d3b35e2cc21d9005e8740fc58675e7a3f589b963ac8cdadd49589d7e19eadf87fd599197788ff4c11d01d98fb809f4eb4b6d60955783da31d9b155d49d9222fcc00c9cb12fc9d266194fcd8a362493a359d87827c5b7a4ac36eb03df3966f942976d4cadf7304047635f9ca4997c09a7Baseline
Run: https://github.com/microsoft/BC-Bench/actions/runs/36851581125; conclusion: success; wall clock: 7.4 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.
synthetic__upgrade-demo-data-not-upgraded-01Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.
Candidate
Run: https://github.com/microsoft/BC-Bench/actions/runs/36852388170; conclusion: success; wall clock: 25.5 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.
synthetic__upgrade-demo-data-not-upgraded-01Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.
Missing telemetry is unavailable, not zero. Evaluation-only metrics exclude candidate generation and are not the full-cycle cost.
Human review must verify source-patch fidelity, gold correctness, and target attribution. No automatic merge or branch-protection claim is made.
Cycle usage and elapsed time (running, as of 2026-10-01T11:23:46.6762740+00:00)
Orchestrator work: 2433.978 seconds. This is not the completed GitHub workflow duration.
Subtotals are CLI/Bench-reported usage, not invoices. Missing values are not zero. Emitted-span coverage does not prove complete child/background billing; judge usage is not exported. Premium requests are not AI credits or currency.
Nested stage, agent and API durations are not added into cycle wall clock. Root queue/setup, telemetry export, artifact upload and later dashboard publication are excluded.
PR evidence is an as-of snapshot before publication completes. Final-for-orchestrator metrics are emitted to self-improvement-cycle.json in the root self-improvement-feedback artifact and its Actions summary when persistence succeeds; a running checkpoint is not final evidence.
Root run: https://github.com/microsoft/BC-ALAgentsInternal/actions/runs/36850344542/attempts/2