Hi there! 👋 I'm Shengxuming Zhang (张圣旭明). ✨
- 🎓 Ph.D. candidate in Software Engineering at Zhejiang University (expected Dec 2026), advised by Prof. Zunlei Feng and Prof. Mingli Song.
- 🔭 My research asks how multimodal foundation models can reason faithfully over visual evidence while staying efficient enough to deploy, through RL post-training of VLMs (verifiable rewards, counterfactual credit assignment, on-policy distillation), unified multimodal LLMs that ground every output in the image, and efficient gigapixel inference, with computational pathology as a testbed.
- 📄 First-author papers at ACM MM 2026 (Oral), CVPR 2023, ECCV 2024 and ICIP 2025. See my Google Scholar.
- 🛠️ Creator of awesome-perception-aware-rlvr, a verl-based framework implementing 18 perception-aware RLVR and on-policy distillation methods, and a contributor to ms-swift.
I'm on the job market for full-time positions in multimodal LLMs and post-training. Feel free to reach me at zsxm1998@qq.com :)

