Zong, Zhuofan
- Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
2024/03/25 by Hao Shao, Shengju Qian, Shao, Hao +13 · 128 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Topic Modeling
- DETRs with Collaborative Hybrid Assignments Training
2022/11/22 by Zhuofan Zong, Zong, Zhuofan, Guanglu Song +3 · 35 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning and Data Classification #Music and Audio Processing
- RAPHAEL: Text-to-Image Generation via Large Mixture of Diffusion Paths
2023/05/29 by Zeyue Xue, Xue, Zeyue, Guanglu Song +11 · 16 citations
Computer Science · Social Sciences · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Computational and Text Analysis Methods
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
2025/05/01 by Jiang, Dongzhi, Guo, Ziyu, Zhang, Renrui +6 · 49 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- CoMat: Aligning Text-to-Image Diffusion Model with Image-to-Text Concept Matching
2024/04/04 by Jiang, Dongzhi, Song, Guanglu, Wu, Xiaoshi +5 · 16 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- MoVA: Adapting Mixture of Vision Experts to Multimodal Context
2024/04/19 by Zhuofan Zong, Zong, Zhuofan, Bingqi Ma +13 · 16 citations
Social Sciences · Engineering · #Geographic Information Systems Studies #Spatial Cognition and Navigation
- Temporal Enhanced Training of Multi-view 3D Object Detector via Historical Object Prediction
2023/04/03 by Zhuofan Zong, Dongzhi Jiang, Zong, Zhuofan +11 · 8 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
- Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models
2024/06/17 by Bingqi Ma, Ma, Bingqi, Zhuofan Zong +7 · 8 citations
Computer Science · #Topic Modeling
- EasyRef: Omni-Generalized Group Image Reference for Diffusion Models via Multimodal LLM
2024/12/12 by Zhuofan Zong, Zong, Zhuofan, Dongzhi Jiang +13 · 8 citations
Medicine · #Radiomics and Machine Learning in Medical Imaging
- Self-slimmed Vision Transformer
2021/11/24 by Zhuofan Zong, Zong, Zhuofan, Kunchang Li +11 · 2 citations
Computer Science · Engineering · #Advanced Neural Network Applications #CCD and CMOS Imaging Sensors #Visual Attention and Saliency Detection