vix.ing · top · new · best · stats · spec

From Representation to Reasoning: Towards both Evidence and Commonsense Reasoning for Video Question-Answering

2022/05/30 by Jiangtong Li, Li Niu, Li, Jiangtong +3 · 3 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Topic Modeling

paper · pdf · doi:10.48550/arxiv.2205.14895

openalex publication_date 2022/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Video understanding has achieved great success in representation learning, such as video caption, video object grounding, and video descriptive question-answer. However, current methods still struggle on video reasoning, including evidence reasoning and commonsense reasoning. To facilitate deeper video understanding towards video reasoning, we present the task of Causal-VidQA, which includes four types of questions ranging from scene description (description) to evidence reasoning (explanation) and commonsense reasoning (prediction and counterfactual). For commonsense reasoning, we set up a two-step solution by answering the question and providing a proper reason. Through extensive experiments on existing VideoQA methods, we find that the state-of-the-art methods are strong in descriptions but weak in reasoning. We hope that Causal-VidQA can guide the research of video understanding from representation learning to deeper reasoning. The dataset and related resources are available at \urlhttps://github.com/bcmi/Causal-VidQA.git.

Cited by

Related