Yue, Zhengrong
- TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
2024/10/25 by X. Zeng, Kunchang Li, Zeng, Xiangyu +22 · 37 citations
Medicine · Computer Science · #COVID-19 diagnosis using AI #Speech Recognition and Synthesis #Video Analysis and Summarization
- VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
2025/06/06 by Zikang Wang, Wang, Zikang, Boyu Chen +11 · 14 citations
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
- LVAgent: Long Video Understanding by Multi-Round Dynamical Collaboration of MLLM Agents
2025/03/13 by Chen, Boyu, Yue, Zhengrong, Chen, Siran +4 · 11 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
2025/09/25 by Yan, Ziang, Li, Xinhao, He, Yinan +6 · 13 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
2025/03/15 by Yue, Zhengrong, Zhuang, Shaobin, Li, Kunchang +2 · 5 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
2025/10/12 by Yue, Zhengrong, Zhang, Haiyu, Zeng, Xiangyu +8 · 5 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences