fix(build): honor CLI optimization bypass - #1401
Merged
Qiong Wu (qiowu) (DingmaomaoBJTU) merged 3 commits intoSep 29, 2026
Merged
Qiong Wu (qiowu) (DingmaomaoBJTU) merged 3 commits into
Qiong Wu (qiowu) (DingmaomaoBJTU) merged 3 commits into
Conversation
Qiong Wu (qiowu) (DingmaomaoBJTU)
marked this pull request as ready for review
September 29, 2026 06:45
Qiong Wu (qiowu) (DingmaomaoBJTU)
requested a review
from a team
as a code owner
September 29, 2026 06:45
…cli-skip-optimize
Qiong Wu (qiowu) (DingmaomaoBJTU)
marked this pull request as draft
September 29, 2026 07:19
Qiong Wu (qiowu) (DingmaomaoBJTU)
marked this pull request as ready for review
September 29, 2026 07:22
xieofxie
approved these changes
Sep 29, 2026
Qiong Wu (qiowu) (DingmaomaoBJTU)
merged commit Sep 29, 2026
0602891
into
microsoft:main
9 checks passed
Qiong Wu (qiowu) (DingmaomaoBJTU)
added a commit
that referenced
this pull request
Sep 29, 2026
## Summary Add an optimization-enabled FP32 QNN/NPU recipe for usyd-community/vitpose-plus-base (keypoint-detection). Only FP32 is included. The unaccepted W8A8 target candidate has been removed; existing generic recipes are unchanged. The model consumes one normalized RGB person crop [1,3,256,192] and emits 17 keypoint heatmaps [1,17,64,48]. Person detection is external. The recipe retains eager attention, GELU/MatMulAdd fusion, and an external QNN EP context. ## Validation Validated on main 0602891 plus the recipe, using model revision 92be54d7a29e42fad47b6e2ca01dd9e685a61e0d and QNN catalog EP 2.2480.53.0. The final edit changes only the advisory note/JSON formatting and removes W8A8; FP32 execution settings are unchanged. - Recipe-authoritative build with -c, -m and -o only: success, 90.0 s (export 44.4 s, optimize 15.8 s, compile 15.4 s). - Short perf smoke: 3 warmups, 10 measured iterations; mean 9.46 ms, p50 9.35 ms, 105.73 samples/s. Not a controlled performance benchmark. - COCO subset: 10 images, all 26 annotated persons. FP32 QNN mAP 0.6761, matching CPU baseline mAP 0.6761 at displayed precision. - Identical real preprocessed crops, QNN vs CPU baseline: cosine mean 1.0000 (rounded), minimum 0.9997; max absolute difference 0.0256. - Baseline built with --no-optimize --no-quant --no-compile after #1401. The short Optimize label is the copy stage: decoded export/optimized GraphProto and opset imports are exactly equal (699 nodes each), despite different serialized file hashes. - 26 recipe-discovery tests passed after narrowing scope; git diff --check passed. These are bounded build/runtime/fidelity checks, not representative COCO accuracy or full model-family support. No FP16, W8A8 or W8A16 acceptance is claimed. Model/processor IDs in the recipe are not revision-pinned. Target-specific discovery selects only FP32 for qnn/npu, rather than unioning generic precision recipes. ## Reproduce ```powershell python -m winml.modelkit build -c examples/recipes/usyd-community_vitpose-plus-base/qnn/npu/keypoint-detection_fp32_config.json -m usyd-community/vitpose-plus-base -o out/vitpose-fp32 python -m winml.modelkit perf -m out/vitpose-fp32/model.onnx --ep qnn --device npu --warmup 3 --iterations 10 python scripts/build_coco_keypoints.py --output-dir out/coco-keypoints --num-images 10 python -m winml.modelkit eval -m out/vitpose-fp32/model.onnx --model-id usyd-community/vitpose-plus-base --task keypoint-detection --dataset out/coco-keypoints --samples 10 --no-shuffle --ep qnn --device npu ```
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fix CLI optimization bypass for single-model Hugging Face and direct-ONNX builds. Both CLI sinks now pass
CLI skip OR config.skip_optimizeto the existing_run_optimize_stagehelper. This also covers CLI-flag fan-out to composite HF components.The shared helper preserves intermediate copying and config behavior while suppressing optimization and its Analyze/autoconf re-optimization loop. Configured FP16/static quantization, compilation, and finalization continue. Ordinary bypass does not mark a raw model prequantized or clear its quantization configuration; QDQ/QOperator detection remains graph-driven.
Only two files differ from main:
commands/build.py(+4/-1) and its regression tests (+251). No recipes, dependencies, new optimization passes, or performance claims.Validation
Updated on 2026-09-29 by merging current main
26e58eadaf396ed5200da158969e125cc33e7406into PR headbd08c0df729f6ccb203908d9702ff90b86114cc4. The merge was conflict-free and retains current-main CGC handling.Fresh local checks on that exact merged head, using Python 3.11.16 and the existing environment:
pytest tests/unit/commands/test_build.py tests/unit/build tests/unit/commands/test_cgc_target_args.py -q: 266 passed, no failures or skips; one upstream Torch deprecation warning.git diff --check upstream/main...HEADpassed.Tests use actual Click dispatch and stage helpers with pytest-generated raw/QDQ/QOperator graphs; expensive export, optimization, quantization, compilation, and host/UI boundaries are mocked. GitHub checks rerun on the new pushed head; their live status is authoritative.
Scope and integration
This validates control flow, not real-model export/calibration/inference, numerical quality, hardware support, or speedup. No full-repository local suite is claimed. Array/module CLI mode, outer composite config-only inheritance, and optimized-bundle generic controls remain outside this fix.
The CLI repair overlaps open #1370 by ssss141414. This PR retains the existing optimize-stage helper contract; #1370 bypasses that helper and also includes model recipes. Coordinate which repair lands and adapt the other PR rather than merging both unchanged. No claim of agreement from that author is made.
The PR is currently ready for review (not draft). This update does not constitute approval or merge; independent review and overlap coordination remain outstanding.