Michele Cafagna
- ViLMA: A Zero-Shot Benchmark for Linguistic and Temporal Grounding in Video-Language Models
2023/11/13 by İlker Kesen, Kesen, Ilker, Andrea Pedrotti +19 · 4 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media