My primary research focuses on the area of vision and language. Presently, I delve into the application of large language models (LLMs) across various tasks involving both vision and language, e.g., language-driven video understanding and open-vocabulary multi-label image recognition. Prior to this, my work revolved around hand detection, hand pose estimation, face recognition, and person re-identification.
My publications can be found at Google Scholar .
- Personal Pages: https://shuoyang129.github.io
- Google Scholar: https://scholar.google.com/citations?user=JJEEfUIAAAAJ
- ORCID: https://orcid.org/0000-0003-2868-7070
- DBLP: https://dblp.org/pid/78/1102-2.html
- 2026.07 🎉🎉 An open-vocabulary multi-label action recognition paper is accepted by CVIU 2026 (CCF-B, JCR Q2, IF=3.6)!
- 2026.05 🎉🎉 An interactive 3D grounding framework and dataset paper is accepted by ICML 2026 (CCF-A conference)!
- 2025.12 🎉🎉 An image-free multi-label image recognition paper is accepted by Pattern Recognition 2026 (中科院一区, JCR Q1, IF=7.6)!
- 2025.06 🎉🎉 An image-text matching paper is accepted by ICCV 2025 (CCF-A conference)!
- 2025.04 🎉🎉 A video visual relationship detection paper is accepted by IJCAI 2025 (CCF-A conference)!
- 2025.04 🎉🎉 A video visual relationship detection paper is accepted by IEEE TPAMI 2025 (CCF-A, 中科院一区, JCR Q1, IF=20.8)!
- 2025.01 🎉🎉 An open-vocabulary multi-label action classification paper is published in 《计算机研究与发展》 2025 (CCF-A Chinese, IF=2.65)!
- Earlier news See the complete news archive on my personal website.
