Guan, Yushuo
- VidCapBench: A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
2025/02/18 by Xinlong Chen, Yuanxing Zhang, Chen, Xinlong +17 · 9 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Video Analysis and Summarization
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
2025/04/14 by Yang Shi, Jiaheng Liu, Shi, Yang +26 · 10 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
2025/06/10 by Chen, Xinlong, Zhang, Yuanxing, Guan, Yushuo +9 · 8 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
2025/05/27 by Yang Shi, Huanqian Wang, Shi, Yang +43 · 5 citations
Computer Science · #Natural Language Processing Techniques #Video Analysis and Summarization #Speech and dialogue systems