vix.ing · top · new · best · stats · spec

Enhancing Video-Language Representations With Structural Spatio-Temporal Alignment

2024/12/01 by Hao Fei, Shengqiong Wu, Meishan Zhang +3 · 6 citations

paper · doi:10.1109/tpami.2024.3393452

Cited by