CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
2025/05/22 by Shilin Yan, Jiaming Han, Yan, Shilin +12 · 7 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.2505.17020
openalex publication_date 2025/05/22 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
The advent of Large Multimodal Models (LMMs) has significantly enhanced Large Language Models (LLMs) to process and interpret diverse data modalities (e.g., image and video). However, as input complexity increases, particularly with long video sequences, the number of required tokens has grown significantly, leading to quadratically computational costs. This has made the efficient compression of video tokens in LMMs, while maintaining performance integrity, a pressing research challenge. In this paper, we introduce CrossLMM, decoupling long video sequences from LMMs via a dual cross-attention mechanism, which substantially reduces visual token quantity with minimal performance degradation. Specifically, we first implement a significant token reduction from pretrained visual encoders through a pooling methodology. Then, within LLM layers, we employ a visual-to-visual cross-attention mechanism, wherein the pooled visual tokens function as queries against the original visual token set. This module enables more efficient token utilization while retaining fine-grained informational fidelity. In addition, we introduce a text-to-visual cross-attention mechanism, for which the text tokens are enhanced through interaction with the original visual tokens, enriching the visual comprehension of the text tokens. Comprehensive empirical evaluation demonstrates that our approach achieves comparable or superior performance across diverse video-based LMM benchmarks, despite utilizing substantially fewer computational resources.
Citations
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
- SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- DynaPrompt: Dynamic Test-Time Prompt Tuning
- LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models
- Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
- VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
- Qwen2.5 Technical Report
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
- AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning
- Accelerating Multimodal Large Language Models by Searching Optimal Vision Token Reduction
- Efficient Multi-modal Large Language Models via Visual Token Grouping
- MC-LLaVA: Multi-Concept Personalized Vision-Language Model
- MC-LLaVA: Multi-Concept Personalized Vision-Language Model
- LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
- AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
- MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
- Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
- LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
- Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
- MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?
- EE-MLLM: A Data-Efficient and Compute-Efficient Multimodal Large Language Model
- LongVILA: Scaling Long-Context Visual Language Models for Long Videos
- LLaVA-OneVision: Easy Visual Task Transfer
- The Llama 3 Herd of Models
- LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
- MAVIS: Mathematical Visual Instruction Tuning with an Automatic Data Engine
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
- A Sanity Check for AI-generated Image Detection
- Long Context Transfer from Language to Vision
- VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
- VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
- Vript: A Video Is Worth Thousands of Words
- Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
- DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
- Visual Perception by Large Language Model's Weights
- LLM as Dataset Analyst: Subpopulation Structure Discovery with Large Language Model
- MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
- How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
- Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
- Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want
- LLaVA-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models
- MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?
- OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning
- Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
- ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
- InstructSeq: Unifying Vision Tasks with Instruction-conditioned Multi-modal Sequence Generation
- LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
- MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
- SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models
- CapsFusion: Rethinking Image-Text Data at Scale
- PanoVOS: Bridging Non-panoramic and Panoramic Views with Transformer for Video Segmentation
- Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
- MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
- MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
- Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation
- Perception Test: A Diagnostic Benchmark for Multimodal Video Models
- LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
- Visual Instruction Tuning
- Sigmoid Loss for Language Image Pre-Training
- GPT-4 Technical Report
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- Video Swin Transformer
- Video Swin Transformer
- Gaussian Error Linear Units (GELUs)
- InternVideo2: Scaling Foundation Models for Multimodal Video Understanding
- MLVU: Benchmarking Multi-task Long Video Understanding
- LLaVA-Video: Video Instruction Tuning With Synthetic Data
Cited by
Related