vix.ing · top · new · best · stats · spec

Boshen Xu

  1. Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
    2025/03/17 by Ye Wang, Ziheng Wang, Wang, Ye +30 · 49 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Domain Adaptation and Few-Shot Learning
  2. Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions?
    2024/05/28 by Boshen Xu, Ziheng Wang, Xu, Boshen +9 · 1 citation
    Psychology · Social Sciences · #Action Observation and Synchronization #Social and Intergroup Psychology