Skip to content

Add Cua-S1 multimodal reference tensor exports - #53

Merged
hsliuustc0106 merged 2 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-multimodal-fixtures
Oct 4, 2026
Merged

hsliuustc0106 merged 2 commits into
ThinkFlowLab:mainfrom
Levius-Fubuki:codex/cua-multimodal-fixtures

Conversation

@Levius-Fubuki

@Levius-Fubuki Levius-Fubuki commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

Export reproducible multimodal boundary tensors from the correctness reference merged in #12, for the native vision integration discussed in #10.

Only two core implementation files are in this PR:

  • recipe/cua_s1/export_multimodal_reference.py: generate seven synthetic requests and observe the actual unmerged PEFT forward. Export preprocessing, token ids, adapted vision features, language inputs, 3D positions, final hidden state and candidate readout as safetensors. A manifest records shapes/dtypes, tensor/file hashes, pinned artifacts and source/execution provenance.
  • recipe/cua_s1/verify_multimodal_reference.py: check bundle integrity, tensor relations, answers, same-image features and exact equality of independent exports.

No tests, documentation, CI changes, dependency lists or generated reports are included in the diff. They were used for validation and retained outside this PR. Serving is unchanged, and this does not depend on #50 or implement native vision execution. #19 is now merged; these boundary tensors support its subsequent language-input integration.

For existing upstream-verified weights, with weights.lock.json next to Qwen3.5-4B/, use a Python 3.12 reference environment matching PACKAGES in the exporter, without FLA/causal-conv1d or explicitly enabled hub kernels:

PYTHONPATH=src HF_HUB_OFFLINE=1 TOKENIZERS_PARALLELISM=false \
  python recipe/cua_s1/export_multimodal_reference.py \
  --weights /path/to/weights --output /tmp/cua-reference-a
# Run a second independent invocation with --output /tmp/cua-reference-b.
PYTHONPATH=src python recipe/cua_s1/verify_multimodal_reference.py \
  /tmp/cua-reference-a --compare /tmp/cua-reference-b

Each output must be a new directory. The manifest is written last, after every captured readout matches an ordinary #12 score call. Temporary hooks are removed on success and failure. Each question produces 14 tensors, including image_features [I,2560], inputs_embeds [1,S,2560] and position_ids [3,1,S]. The base, processor and adapter configurations accompany each generated bundle; raw BF16 bytes are preserved.

Test Plan

  • Local validation suite, retained outside the PR: CUDA_VISIBLE_DEVICES="" PYTHONPATH=src python -m pytest tests/cua_s1/test_export_multimodal_reference.py -q.
  • ruff check --isolated --select E4,E7,E9,F,I recipe/cua_s1/export_multimodal_reference.py recipe/cua_s1/verify_multimodal_reference.py.
  • ruff format --isolated --check on those two files, git diff --check and CLI help.
  • Two independent exports and verify_multimodal_reference.py <run-a> --compare <run-b>.

System1-Omni Version / Commit: 1b64fa2, based on main b50aa28. The two implementation files are byte-identical to the validated source.

Test Result

  • 11 tests passed with CUDA devices hidden, without checkpoints or a serving process. They cover deterministic inputs, output protection, BF16 byte fingerprints, actual hook capture/cleanup, safetensors round trips, file corruption and inventory/path checks, missing tensors and CPU softmax roundoff versus drift.
  • Lint, formatting, diff and CLI checks passed.
  • On one RTX 4090 (24 GiB), driver 595.71.05, Torch 2.14.0+cu130, Transformers 5.17.0 and PEFT 0.21.0: each export contains 8 questions, 112 tensors and 25 payload files. Both independent manifests and all raw tensor hashes match exactly. Every captured probability matches a separate ordinary Add Cua-S1 0.2 multimodal CUDA worker with upstream parity #12 score() call exactly.
  • Inputs span five image sizes, PNG/JPEG, 1/3/26 candidates and a same-image two-question request with structured/non-ASCII text. Shape/dtype/finite-value, patch/grid/merge, placeholder ordering, image insertion, rope delta, candidate ordering, answer reconstruction and same-image feature checks pass.
  • CPU softmax reconstruction differs from the GPU reference by at most 3.73e-9; that formula check permits absolute error 1e-7 with zero relative tolerance. Repeat exports still require exact equality. Bitwise repeatability has only been checked on the recorded GPU/software setup.
  • No native accuracy/performance, FP32 control, other GPU, video, batching or Metal claim. Rust source is unchanged; Rust workspace checks are left to the existing CI job.

Self-review

Agent-assisted full-diff review found no blocking issue. The full implementation and archived validation were rechecked on 2026-10-01: the 11 CPU tests and two independent bundle verification pass again. Ready for maintainer review; agent review does not represent human contributor sign-off or maintainer approval.

  • I have reviewed the full diff and addressed the issues I found.
  • I have checked that the change follows the project's architecture and stays focused on the stated purpose.
  • I have run the checks appropriate to this change and reported commands, results, and anything I could not verify above.
  • I have checked that the PR description, documentation, and any accuracy or performance claims match the implementation and available evidence.

@Levius-Fubuki
Levius-Fubuki marked this pull request as ready for review October 1, 2026 01:49
@Levius-Fubuki
Levius-Fubuki requested review from hsliuustc0106 and a balanced review from Copilot and removed request for Copilot October 1, 2026 01:49
@Levius-Fubuki

Copy link
Copy Markdown
Collaborator Author

@hsliuustc0106 @twu3202 #53 is ready for review. Its two implementation files remain at 1b64fa2, byte-identical to the GPU-validated source. I re-ran the 11 archived CPU tests and verification of the two independent RTX 4090 exports today: 8 questions, 112 tensors and 25 payload files, with exact independent-export equality. Lint/format/diff checks pass.

I am implementing the subsequent native language boundary discussed in #10: token IDs plus adapted BF16 image features in placeholder order, and explicit int64 T/H/W positions corresponding to [3,1,S]. It will reuse #19's language loop and CUDA kernels, preserve the text HTTP worker, and use this PR's fixed tensors for numerical validation. The image processor and vision encoder stay in the reference path for this step. I will link the focused implementation PR and raw validation evidence when ready.

@Levius-Fubuki

Copy link
Copy Markdown
Collaborator Author

@hsliuustc0106 The subsequent native language boundary is ready in #56: token IDs plus BF16 image features and explicit int64 T/H/W positions, using the existing CUDA language loop with interleaved MRoPE. Native vision/preprocessing and image HTTP serving remain outside this step.

Immutable validation evidence now includes the full #53 boundary bundle, the independent-export manifest/equality report, native/control outputs, source hashes, scripts and test logs. It ties this export revision 1b64fa2 to native revision 93a6cb0. I also reran all 11 exporter CPU tests today.

On RTX 4090, all 8 questions (16 repeated native forwards) passed the declared language-boundary numerical gate. Maximum native probability error against the fixed-boundary FP32 language control is 0.00153646, versus an allowance of 0.02472938; all top choices match. These results do not assert bitwise HF parity or end-to-end FP32 vision accuracy. The three native GPU regression/kernel tests also pass, including disjoint image spans and text-position restoration after scratch growth.

Please review #53's captured reference boundary and #56's consumer contract together; the implementation PR keeps raw evidence out of its six-file runtime diff.

@hsliuustc0106
hsliuustc0106 merged commit 6d92519 into ThinkFlowLab:main Oct 4, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants