vix.ing · top · new · best · stats · spec

VideoMamba: Spatio-Temporal Selective State Space Model

2024/07/11 by Jinyoung Park, Park, Jinyoung, Hee-Seon Kim +7 · 5 citations
Physics and Astronomy · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Opinion Dynamics and Social Influence

paper · pdf · doi:10.48550/arxiv.2407.08476

openalex publication_date 2024/07/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

We introduce VideoMamba, a novel adaptation of the pure Mamba architecture, specifically designed for video recognition. Unlike transformers that rely on self-attention mechanisms leading to high computational costs by quadratic complexity, VideoMamba leverages Mamba's linear complexity and selective SSM mechanism for more efficient processing. The proposed Spatio-Temporal Forward and Backward SSM allows the model to effectively capture the complex relationship between non-sequential spatial and sequential temporal information in video. Consequently, VideoMamba is not only resource-efficient but also effective in capturing long-range dependency in videos, demonstrated by competitive performance and outstanding efficiency on a variety of video understanding benchmarks. Our work highlights the potential of VideoMamba as a powerful tool for video understanding, offering a simple yet effective baseline for future research in video analysis.

Cited by

Related