Skip to content

recipe(vitpose): add FP32 QNN NPU recipe with bounded validation - #1419

Merged
Qiong Wu (qiowu) (DingmaomaoBJTU) merged 2 commits into
microsoft:mainfrom
DingmaomaoBJTU:dingmaomaobjtu/add-usyd-community-vitpose-plus-base-recipe
Sep 29, 2026
Merged

Qiong Wu (qiowu) (DingmaomaoBJTU) merged 2 commits into
microsoft:mainfrom
DingmaomaoBJTU:dingmaomaobjtu/add-usyd-community-vitpose-plus-base-recipe

Conversation

@DingmaomaoBJTU

@DingmaomaoBJTU Qiong Wu (qiowu) (DingmaomaoBJTU) commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add an optimization-enabled FP32 QNN/NPU recipe for usyd-community/vitpose-plus-base (keypoint-detection). Only FP32 is included. The unaccepted W8A8 target candidate has been removed; existing generic recipes are unchanged.

The model consumes one normalized RGB person crop [1,3,256,192] and emits 17 keypoint heatmaps [1,17,64,48]. Person detection is external. The recipe retains eager attention, GELU/MatMulAdd fusion, and an external QNN EP context.

Validation

Validated on main 0602891 plus the recipe, using model revision 92be54d7a29e42fad47b6e2ca01dd9e685a61e0d and QNN catalog EP 2.2480.53.0. The final edit changes only the advisory note/JSON formatting and removes W8A8; FP32 execution settings are unchanged.

  • Recipe-authoritative build with -c, -m and -o only: success, 90.0 s (export 44.4 s, optimize 15.8 s, compile 15.4 s).
  • Short perf smoke: 3 warmups, 10 measured iterations; mean 9.46 ms, p50 9.35 ms, 105.73 samples/s. Not a controlled performance benchmark.
  • COCO subset: 10 images, all 26 annotated persons. FP32 QNN mAP 0.6761, matching CPU baseline mAP 0.6761 at displayed precision.
  • Identical real preprocessed crops, QNN vs CPU baseline: cosine mean 1.0000 (rounded), minimum 0.9997; max absolute difference 0.0256.
  • Baseline built with --no-optimize --no-quant --no-compile after fix(build): honor CLI optimization bypass #1401. The short Optimize label is the copy stage: decoded export/optimized GraphProto and opset imports are exactly equal (699 nodes each), despite different serialized file hashes.
  • 26 recipe-discovery tests passed after narrowing scope; git diff --check passed.

These are bounded build/runtime/fidelity checks, not representative COCO accuracy or full model-family support. No FP16, W8A8 or W8A16 acceptance is claimed. Model/processor IDs in the recipe are not revision-pinned. Target-specific discovery selects only FP32 for qnn/npu, rather than unioning generic precision recipes.

Reproduce

python -m winml.modelkit build -c examples/recipes/usyd-community_vitpose-plus-base/qnn/npu/keypoint-detection_fp32_config.json -m usyd-community/vitpose-plus-base -o out/vitpose-fp32
python -m winml.modelkit perf -m out/vitpose-fp32/model.onnx --ep qnn --device npu --warmup 3 --iterations 10
python scripts/build_coco_keypoints.py --output-dir out/coco-keypoints --num-images 10
python -m winml.modelkit eval -m out/vitpose-fp32/model.onnx --model-id usyd-community/vitpose-plus-base --task keypoint-detection --dataset out/coco-keypoints --samples 10 --no-shuffle --ep qnn --device npu

@DingmaomaoBJTU Qiong Wu (qiowu) (DingmaomaoBJTU) added the model-scale-by-skill Model support PR created or maintained by the adding-model-support skill label Sep 15, 2026
@DingmaomaoBJTU

Copy link
Copy Markdown
Collaborator Author

Independent agent limited review, relayed by the publishing agent.

LIMITED_SCOPE_HONESTY_CHECK: NO_FINDINGS

Independent post-publication review under charter revision 2 and the explicit “先发待验证 Draft” exception. This is not model-support approval, Goal acceptance, ordinary workflow approval, a GitHub Review, or permission to merge or mark ready.

Reviewed the actual published body and complete diff at head 1fd43e7, base/parent da5dbcd.

  • Scope and publication: OPEN/Draft, author DingmaomaoBJTU, label model-scale-by-skill; exactly two added QNN/NPU FP32/W8A8 JSON candidates, +57/+77, no deletions or CLI/source/test/index changes. Remote blobs, committed blobs, local bytes, and the tester's frozen SHA256s match.
  • Reporting: the three required headings and seven evidence sections are present. Both first-field _note warnings are candid. The body admits that _note is ignored/advisory, not an execution gate, and exact-target discovery selects FP32/W8A8 instead of unioning generic FP16. Full historical deltas are 0/1; anchor deltas 5/7. The portable processor ID is not revision-pinned; dummy range, seed, and distribution do not prove normalized inputs or calibration provenance. The frozen source-only profile retains its selected-expert caveat, and public processor-path redaction is explicit.
  • Saved evidence, not rerun tests: logs, JUnit identities, and seals corroborate 97 = 34 config + 26 discovery + 9 policy + 28 scratch checks, the 60-test subset replay (not additive), and 8 metadata checks. The earlier bootstrap failure, six harness failures, scratch-only corrections, and native ORT import warning are disclosed. Protected small-text hashes, including the lock and transcript, remain unchanged.
  • Unvalidated boundaries: all current L0–L3 rows for FP32/FP16/W8A8/W8A16, FP32 CPU functional evaluation, and component/operator analysis remain NOT_RUN. No highest Goal, current performance, or support result is established. Historical FP32 minimum cosine 0.9999722983142216 is old numeric evidence only. W8A8 full-set 0.982657340445321 and wallpaper 0.6957623048200067 fail 0.99; posthoc 0.9937934557606259 does not rescue them. W8A16 0.4100480666185648 remains a failure, with no target FP16/W8A16 recipe. Historical 2.5840062576234804x, 8 pairs/16 sessions/20 warmups/100 measurements, is not new or correctness-gated performance. QNN memory remains NOT_MEASURED; the older one-image CPU AP is operability-only, not representative or W8A8 accuracy.
  • Dependency and commands: fix(build): honor CLI optimization bypass #1401 remains a separate canonical raw/no-optimize baseline dependency, not a proven optimization-enabled runtime dependency. Proposed builds are NOT_EXECUTED; the public 60-test command is distinguished from both the guarded checks and model work.

At 2026-09-15 17:09 UTC, all 9 visible GitHub checks were successful; the earlier publication snapshot was pending. Neither state validates these model candidates. Before this comment, there were 0 conversation comments, 0 review threads, and 0 GitHub reviews.

Link-check limitation: seven linked GitHub files and #1401 resolved through live metadata. The two pinned Hugging Face link patterns and existing source-file hashes match the frozen identity; a live Hugging Face metadata request failed at the network connection, so remote reachability is not claimed.

Review verification used read-only Git/GitHub and small text/JSON/XML/hash inspection. No Python/project imports, tests, CLI help, model commands, dependency/provider operations, payload downloads, cleanup, source/recipe edits, or auth changes were performed. No fixes are requested within this limited review; intentional missing model validation is not a finding against the authorized unvalidated Draft. This opinion may be delivered only as a normal comment, never a GitHub Review.

@DingmaomaoBJTU
Qiong Wu (qiowu) (DingmaomaoBJTU) force-pushed the dingmaomaobjtu/add-usyd-community-vitpose-plus-base-recipe branch from 1fd43e7 to b0660d0 Compare September 29, 2026 08:03
@DingmaomaoBJTU Qiong Wu (qiowu) (DingmaomaoBJTU) changed the title recipe(vitpose): add unvalidated QNN NPU candidates recipe(vitpose): add FP32 QNN NPU recipe with bounded validation Sep 29, 2026
@DingmaomaoBJTU
Qiong Wu (qiowu) (DingmaomaoBJTU) marked this pull request as ready for review September 29, 2026 08:04
@DingmaomaoBJTU
Qiong Wu (qiowu) (DingmaomaoBJTU) merged commit 982f7bd into microsoft:main Sep 29, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

model-scale-by-skill Model support PR created or maintained by the adding-model-support skill

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants