Skip to content

fix(build): honor CLI optimization bypass - #1401

Merged
Qiong Wu (qiowu) (DingmaomaoBJTU) merged 3 commits into
microsoft:mainfrom
DingmaomaoBJTU:dingmaomaobjtu/fix-cli-skip-optimize
Sep 29, 2026
Merged

Qiong Wu (qiowu) (DingmaomaoBJTU) merged 3 commits into
microsoft:mainfrom
DingmaomaoBJTU:dingmaomaobjtu/fix-cli-skip-optimize

Conversation

@DingmaomaoBJTU

@DingmaomaoBJTU Qiong Wu (qiowu) (DingmaomaoBJTU) commented Sep 8, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Fix CLI optimization bypass for single-model Hugging Face and direct-ONNX builds. Both CLI sinks now pass CLI skip OR config.skip_optimize to the existing _run_optimize_stage helper. This also covers CLI-flag fan-out to composite HF components.

The shared helper preserves intermediate copying and config behavior while suppressing optimization and its Analyze/autoconf re-optimization loop. Configured FP16/static quantization, compilation, and finalization continue. Ordinary bypass does not mark a raw model prequantized or clear its quantization configuration; QDQ/QOperator detection remains graph-driven.

Only two files differ from main: commands/build.py (+4/-1) and its regression tests (+251). No recipes, dependencies, new optimization passes, or performance claims.

Validation

Updated on 2026-09-29 by merging current main 26e58eadaf396ed5200da158969e125cc33e7406 into PR head bd08c0df729f6ccb203908d9702ff90b86114cc4. The merge was conflict-free and retains current-main CGC handling.

Fresh local checks on that exact merged head, using Python 3.11.16 and the existing environment:

  • pytest tests/unit/commands/test_build.py tests/unit/build tests/unit/commands/test_cgc_target_args.py -q: 266 passed, no failures or skips; one upstream Torch deprecation warning.
  • Changed-file Ruff lint and format checks passed.
  • git diff --check upstream/main...HEAD passed.
  • Before updating the PR, the same 40 new regression cases against current-main production code reproduced 14 failures / 26 passes, confirming the bypass defect still exists.

Tests use actual Click dispatch and stage helpers with pytest-generated raw/QDQ/QOperator graphs; expensive export, optimization, quantization, compilation, and host/UI boundaries are mocked. GitHub checks rerun on the new pushed head; their live status is authoritative.

Scope and integration

This validates control flow, not real-model export/calibration/inference, numerical quality, hardware support, or speedup. No full-repository local suite is claimed. Array/module CLI mode, outer composite config-only inheritance, and optimized-bundle generic controls remain outside this fix.

The CLI repair overlaps open #1370 by ssss141414. This PR retains the existing optimize-stage helper contract; #1370 bypasses that helper and also includes model recipes. Coordinate which repair lands and adapt the other PR rather than merging both unchanged. No claim of agreement from that author is made.

The PR is currently ready for review (not draft). This update does not constitute approval or merge; independent review and overlap coordination remain outstanding.

Comment thread tests/unit/commands/test_build.py Fixed
@DingmaomaoBJTU
Qiong Wu (qiowu) (DingmaomaoBJTU) marked this pull request as ready for review September 29, 2026 07:22
@DingmaomaoBJTU
Qiong Wu (qiowu) (DingmaomaoBJTU) merged commit 0602891 into microsoft:main Sep 29, 2026
9 checks passed
Qiong Wu (qiowu) (DingmaomaoBJTU) added a commit that referenced this pull request Sep 29, 2026
## Summary

Add an optimization-enabled FP32 QNN/NPU recipe for
usyd-community/vitpose-plus-base (keypoint-detection). Only FP32 is
included. The unaccepted W8A8 target candidate has been removed;
existing generic recipes are unchanged.

The model consumes one normalized RGB person crop [1,3,256,192] and
emits 17 keypoint heatmaps [1,17,64,48]. Person detection is external.
The recipe retains eager attention, GELU/MatMulAdd fusion, and an
external QNN EP context.

## Validation

Validated on main 0602891 plus the
recipe, using model revision 92be54d7a29e42fad47b6e2ca01dd9e685a61e0d
and QNN catalog EP 2.2480.53.0. The final edit changes only the advisory
note/JSON formatting and removes W8A8; FP32 execution settings are
unchanged.

- Recipe-authoritative build with -c, -m and -o only: success, 90.0 s
(export 44.4 s, optimize 15.8 s, compile 15.4 s).
- Short perf smoke: 3 warmups, 10 measured iterations; mean 9.46 ms, p50
9.35 ms, 105.73 samples/s. Not a controlled performance benchmark.
- COCO subset: 10 images, all 26 annotated persons. FP32 QNN mAP 0.6761,
matching CPU baseline mAP 0.6761 at displayed precision.
- Identical real preprocessed crops, QNN vs CPU baseline: cosine mean
1.0000 (rounded), minimum 0.9997; max absolute difference 0.0256.
- Baseline built with --no-optimize --no-quant --no-compile after #1401.
The short Optimize label is the copy stage: decoded export/optimized
GraphProto and opset imports are exactly equal (699 nodes each), despite
different serialized file hashes.
- 26 recipe-discovery tests passed after narrowing scope; git diff
--check passed.

These are bounded build/runtime/fidelity checks, not representative COCO
accuracy or full model-family support. No FP16, W8A8 or W8A16 acceptance
is claimed. Model/processor IDs in the recipe are not revision-pinned.
Target-specific discovery selects only FP32 for qnn/npu, rather than
unioning generic precision recipes.

## Reproduce

```powershell
python -m winml.modelkit build -c examples/recipes/usyd-community_vitpose-plus-base/qnn/npu/keypoint-detection_fp32_config.json -m usyd-community/vitpose-plus-base -o out/vitpose-fp32
python -m winml.modelkit perf -m out/vitpose-fp32/model.onnx --ep qnn --device npu --warmup 3 --iterations 10
python scripts/build_coco_keypoints.py --output-dir out/coco-keypoints --num-images 10
python -m winml.modelkit eval -m out/vitpose-fp32/model.onnx --model-id usyd-community/vitpose-plus-base --task keypoint-detection --dataset out/coco-keypoints --samples 10 --no-shuffle --ep qnn --device npu
```
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants