Skip to content

knowledge: improve review precision from BCApps PR 10990-neg-8222ad38e67a1bf4ee3b67d2389a957f0d00c77a feedback - #206

Draft
Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-10990-neg-8222ad38e67a1bf4ee3b67d2389a957f0d00c77a-quality
Draft

Wenjie Fan (gggdttt) wants to merge 1 commit into
mainfrom
self-improvement/bcapps-10990-neg-8222ad38e67a1bf4ee3b67d2389a957f0d00c77a-quality

Conversation

@gggdttt

@gggdttt Wenjie Fan (gggdttt) commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Improves BCQuality knowledge based on maintainer thumbs-down feedback from explicit BCApps PR review runs.

Source feedback

Validation

  • Frontmatter and article structure validation completed.

Generated by the BC-ALAgentsInternal self-improvement workflow.

Offline evaluation: regression

Candidate correctness failed: unexpected or missed findings remain. Existing ignored gold comments retain their neutral scoring semantics.

Preparation attempt 1: validated.

Selection: new. New: synthetic__upgrade-demo-data-not-upgraded-01. Reused (complete payloads): ``.

  • synthetic__upgrade-demo-data-not-upgraded-01 / calibration_context / Add-Expense-VAT-settings-to-Contoso-demo-tool BCApps#10990 (comment): Reproduces the reviewed boundary: a Contoso demo-data module's CreateMasterData() wires a new VAT-rate seeding codeunit that inserts persisted setup records, with no Subtype = Upgrade codeunit to propagate those records to already-provisioned tenants, matching the bot's upgrade-codeunit-subtype and install-code-does-not-run-on-version-upgrade reasoning. Feedback was negative (THUMBS_DOWN) and the author replied 'Not needed here, this is Preview demo data', rejecting applicability to Contoso/demo-data tooling specifically -- a product/process judgment call, not a confirmed absence of the underlying defect pattern. Per the no-invented-clean-negative rule, the expected finding is retained rather than converted into a false_positive_guard; this is deliberately scoped as calibration_context so the disagreement is captured honestly. Limitation: does not prove the bot's generic upgrade-safety guidance extends to demo-data/Preview-tooling modules, and does not resolve whether BC's Contoso demo-data lifecycle exempts such modules from Subtype = Upgrade requirements -- that remains an open, human-reviewed question.
    Coverage references and outcome shapes are validated mechanically. Semantic equivalence, severity calibration and recommendation quality are NOT proved by these checks or a matching F1.

Common engine code: ecf8e31759d6ddd6d78e3a0b7836b40134368009; dataset SHA256: 46A978950BE2F3B98ED647D3AF503448E74879D29E95988C1394EB90DD443EB6.
Model: gpt-5.6-luna; judge: gpt-5.3-codex. Exact D: synthetic__upgrade-demo-data-not-upgraded-01.
Dataset base: e89b01069c079e3e4741587ffd5b4c45eae85143; candidate: ddb215cb28634e8332a2131f8c992833b718a150.

Arm Evaluation commit Engine pin Knowledge pin
baseline d537505e3e236358d3b35e2cc21d9005e8740fc5 8675e7a3f589b963ac8cdadd49589d7e19eadf87 fd599197788ff4c11d01d98fb809f4eb4b6d6095
candidate 5783da31d9b155d49d9222fcc00c9cb12fc9d266 194fcd8a362493a359d87827c5b7a4ac36eb03df 3966f942976d4cadf7304047635f9ca4997c09a7

Baseline

Run: https://github.com/microsoft/BC-Bench/actions/runs/36851581125; conclusion: success; wall clock: 7.4 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__upgrade-demo-data-not-upgraded-01 1 3 1 3 0 unavailable 302.7516946

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.

Aggregate metric Value
total 1
expected_comment_count 1
generated_comment_count 3
matched_comment_count 0
missed_comment_count 1
incorrect_comment_count 3
precision 0
recall 0
f1 0

Candidate

Run: https://github.com/microsoft/BC-Bench/actions/runs/36852388170; conclusion: success; wall clock: 25.5 minutes.
Raw findings and runtime pin proofs are retained in the self-improvement-feedback artifact and the linked evaluation run.

Entry Expected Generated Missed Unexpected F1 AI credits Agent seconds
synthetic__upgrade-demo-data-not-upgraded-01 1 3 1 3 0 unavailable 1353.393822

Observed AI credits: unavailable; coverage 0/1 entries. Missing entries are not extrapolated into a total; reported amounts do not attest full credit coverage.

Aggregate metric Value
total 1
expected_comment_count 1
generated_comment_count 3
matched_comment_count 0
missed_comment_count 1
incorrect_comment_count 3
precision 0
recall 0
f1 0

Missing telemetry is unavailable, not zero. Evaluation-only metrics exclude candidate generation and are not the full-cycle cost.
Human review must verify source-patch fidelity, gold correctness, and target attribution. No automatic merge or branch-protection claim is made.

Cycle usage and elapsed time (running, as of 2026-10-01T11:23:46.6762740+00:00)

Orchestrator work: 2433.978 seconds. This is not the completed GitHub workflow duration.

Stage / attempt Status / outcome Elapsed seconds
preflight / 1 completed / 5.962
quality_preparation / 1 completed / 92.567
gold_preparation / 1 completed / selected_new 279.745
evaluation_preparation / 1 completed / 29.659
baseline / 1 completed / success 477.288
candidate / 1 completed / success 1547.885
gate_verdict / 1 completed / regression 0.018
publication / 1 running / 0.033
receipts / 1 not_started / unavailable
quality_validation / 1 completed / 3.885
gold_attempt / 1 completed / validated 260.725
Measure Known subtotal Whole-cycle total Known / required components
ai_credit: aiCredits 91.18581 unavailable 2 / 6
premium_request: premiumRequests 53 unavailable 2 / 6
token: inputTokens 14351380 unavailable 4 / 6
token: outputTokens 185600 unavailable 4 / 6
token: totalTokens 14536980 unavailable 4 / 6

Subtotals are CLI/Bench-reported usage, not invoices. Missing values are not zero. Emitted-span coverage does not prove complete child/background billing; judge usage is not exported. Premium requests are not AI credits or currency.
Nested stage, agent and API durations are not added into cycle wall clock. Root queue/setup, telemetry export, artifact upload and later dashboard publication are excluded.
PR evidence is an as-of snapshot before publication completes. Final-for-orchestrator metrics are emitted to self-improvement-cycle.json in the root self-improvement-feedback artifact and its Actions summary when persistence succeeds; a running checkpoint is not final evidence.
Root run: https://github.com/microsoft/BC-ALAgentsInternal/actions/runs/36850344542/attempts/2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed draft head 3966f942976d4cadf7304047635f9ca4997c09a7. The exception is currently justified by a false regeneration model: the article says Contoso/demo modules are explicitly rerun and data is regenerated on demand, but Contoso Demo Tool records generated modules via Data Level, excludes modules for which IsModuleGenerated(...) is true, and errors when selected modules were already generated. Generation also occurs through new-company/build entry points, so it is not exclusively manual. This could suppress legitimate upgrade findings based on a nonexistent rerun path.

Please scope the exception to data explicitly owned by a disposable/unsupported demo-data module where compatibility for previously generated demo companies is intentionally not required. Remove claims that modules are rerun or regenerated on demand, and do not treat generic preview/trial status as proof that persisted data needs no upgrade path. Add a negative evaluation fixture if this suppression must remain deterministic.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants