2024/07/10 by Zhiting Wang, Wang, Zhiting, Qiangong Zhou +7 · 1 citation
Computer Science · Engineering · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Image and Video Processing (eess.IV) #Machine Learning and Data Classification #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Robotics and Automated Systems #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2407.07325
openalex publication_date 2024/07/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This technical report presents the implementation of a state-of-the-art video encoder for video-text modal alignment and a video conversation framework called HiLight, which features dual visual towers. The work is divided into two main parts: 1.alignment of video and text modalities; 2.convenient and efficient way to interact with users. Our goal is to address the task of video comprehension in the context of billiards. The report includes a discussion of the concepts and the final solution developed during the task's implementation.