Add Cua-S1 multimodal reference tensor exports - #53
hsliuustc0106 merged 2 commits into
Conversation
|
@hsliuustc0106 @twu3202 #53 is ready for review. Its two implementation files remain at I am implementing the subsequent native language boundary discussed in #10: token IDs plus adapted BF16 image features in placeholder order, and explicit int64 T/H/W positions corresponding to |
|
@hsliuustc0106 The subsequent native language boundary is ready in #56: token IDs plus BF16 image features and explicit int64 T/H/W positions, using the existing CUDA language loop with interleaved MRoPE. Native vision/preprocessing and image HTTP serving remain outside this step. Immutable validation evidence now includes the full #53 boundary bundle, the independent-export manifest/equality report, native/control outputs, source hashes, scripts and test logs. It ties this export revision On RTX 4090, all 8 questions (16 repeated native forwards) passed the declared language-boundary numerical gate. Maximum native probability error against the fixed-boundary FP32 language control is 0.00153646, versus an allowance of 0.02472938; all top choices match. These results do not assert bitwise HF parity or end-to-end FP32 vision accuracy. The three native GPU regression/kernel tests also pass, including disjoint image spans and text-position restoration after scratch growth. Please review #53's captured reference boundary and #56's consumer contract together; the implementation PR keeps raw evidence out of its six-file runtime diff. |
Purpose
Export reproducible multimodal boundary tensors from the correctness reference merged in #12, for the native vision integration discussed in #10.
Only two core implementation files are in this PR:
recipe/cua_s1/export_multimodal_reference.py: generate seven synthetic requests and observe the actual unmerged PEFT forward. Export preprocessing, token ids, adapted vision features, language inputs, 3D positions, final hidden state and candidate readout as safetensors. A manifest records shapes/dtypes, tensor/file hashes, pinned artifacts and source/execution provenance.recipe/cua_s1/verify_multimodal_reference.py: check bundle integrity, tensor relations, answers, same-image features and exact equality of independent exports.No tests, documentation, CI changes, dependency lists or generated reports are included in the diff. They were used for validation and retained outside this PR. Serving is unchanged, and this does not depend on #50 or implement native vision execution. #19 is now merged; these boundary tensors support its subsequent language-input integration.
For existing upstream-verified weights, with
weights.lock.jsonnext toQwen3.5-4B/, use a Python 3.12 reference environment matchingPACKAGESin the exporter, without FLA/causal-conv1d or explicitly enabled hub kernels:PYTHONPATH=src HF_HUB_OFFLINE=1 TOKENIZERS_PARALLELISM=false \ python recipe/cua_s1/export_multimodal_reference.py \ --weights /path/to/weights --output /tmp/cua-reference-a # Run a second independent invocation with --output /tmp/cua-reference-b. PYTHONPATH=src python recipe/cua_s1/verify_multimodal_reference.py \ /tmp/cua-reference-a --compare /tmp/cua-reference-bEach output must be a new directory. The manifest is written last, after every captured readout matches an ordinary #12 score call. Temporary hooks are removed on success and failure. Each question produces 14 tensors, including
image_features [I,2560],inputs_embeds [1,S,2560]andposition_ids [3,1,S]. The base, processor and adapter configurations accompany each generated bundle; raw BF16 bytes are preserved.Test Plan
CUDA_VISIBLE_DEVICES="" PYTHONPATH=src python -m pytest tests/cua_s1/test_export_multimodal_reference.py -q.ruff check --isolated --select E4,E7,E9,F,I recipe/cua_s1/export_multimodal_reference.py recipe/cua_s1/verify_multimodal_reference.py.ruff format --isolated --checkon those two files,git diff --checkand CLI help.verify_multimodal_reference.py <run-a> --compare <run-b>.System1-Omni Version / Commit:
1b64fa2, based on mainb50aa28. The two implementation files are byte-identical to the validated source.Test Result
score()call exactly.Self-review
Agent-assisted full-diff review found no blocking issue. The full implementation and archived validation were rechecked on 2026-10-01: the 11 CPU tests and two independent bundle verification pass again. Ready for maintainer review; agent review does not represent human contributor sign-off or maintainer approval.