Boshen Xu
- Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
2025/03/17 by Ye Wang, Ziheng Wang, Wang, Ye +30 · 49 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Domain Adaptation and Few-Shot Learning
- Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
2024/05/28 by Boshen Xu, Ziheng Wang, Xu, Boshen +9 · 1 citation
Psychology · Social Sciences · #Action Observation and Synchronization #Social and Intergroup Psychology