Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
2023/06/08 by Muhammad Maaz, Maaz, Muhammad, Hanoona Rasheed +5 · 1 voice · 309 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Topic Modeling #cs.CV
paper · pdf · doi:10.48550/arxiv.2306.05424
openalex publication_date 2023/06/08 · arxiv published 2023/06/08 · openalex created_date 2023/06/10 · arxiv updated 2024/06/10 · openalex updated_date 2026/07/28
Abstract
Conversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data. While there have been initial attempts for image-based conversation models, this work addresses the under-explored field of video-based conversation by introducing Video-ChatGPT. It is a multimodal model that merges a video-adapted visual encoder with an LLM. The resulting model is capable of understanding and generating detailed conversations about videos. We introduce a new dataset of 100,000 video-instruction pairs used to train Video-ChatGPT acquired via manual and semi-automated pipeline that is easily scalable and robust to label noise. We also develop a quantitative evaluation framework for video-based dialogue models to objectively analyze the strengths and weaknesses of video-based dialogue models. Code: https://github.com/mbzuai-oryx/Video-ChatGPT.
Cited by
- TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision
- Video Understanding: From Geometry and Semantics to Unified Models
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia
- Unifying Learning Dynamics and Generalization in Transformers Scaling Law
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
- M3-Verse: A "Spot the Difference" Challenge for Large Multimodal Models
- Enhancing AIGC Service Efficiency with Adaptive Multi-Edge Collaboration in A Distributed System
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- SNOW: Spatio-Temporal Scene Understanding with World Knowledge for Open-World Embodied Reasoning
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Point to Span: Zero-Shot Moment Retrieval for Navigating Unseen Hour-Long Videos
- ChronusOmni: Improving Time Awareness of Omni Large Language Models
- Rethinking Chain-of-Thought Reasoning for Videos
- Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- What Happens When: Learning Temporal Orders of Events in Videos
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- UniComp: Rethinking Video Compression Through Informational Uniqueness
- Divide, then Ground: Adapting Frame Selection to Query Types for Long-Form Video Understanding
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
- Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
- SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
- Learning Skill-Attributes for Transferable Assessment in Video
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- Language-Guided Graph Representation Learning for Video Summarization
- UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
- VidEmo: Affective-Tree Reasoning for Emotion-Centric Video Foundation Models
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Semantic Frame Aggregation-based Transformer for Live Video Comment Generation
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- A Video Is Not Worth a Thousand Words
- Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
- SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- FeatureFool: Zero-Query Fooling of Video Models via Feature Map
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
- Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
- RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- Temporal Prompting Matters: Rethinking Referring Video Object Segmentation
- Addressing the ID-Matching Challenge in Long Video Captioning
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- FrameOracle: Learning What to See and How Much to See in Videos
- AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
- POVQA: Preference-Optimized Video Question Answering with Rationales for Data Efficiency
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- NeMo: Needle in a Montage for Video-Language Understanding
- VideoScore2: Think before You Score in Generative Video Evaluation
- Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
- RefCaptioner: Multi-Reference Image-Grounded Video Captioning
- VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- iFinder: Structured Zero-Shot Vision-Based LLM Grounding for Dash-Cam Video Reasoning
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
- Dense Video Understanding with Gated Residual Tokenization
- UTI-LLM: A Personalized Articulatory-Speech Therapy Assistance System Based on Multimodal Large Language Model
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Open-ended Hierarchical Streaming Video Understanding with Vision Language Models
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- Video Understanding by Design: How Datasets Shape Architectures and Insights
- AdsQA: Towards Advertisement Video Understanding
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Human Motion Video Generation: A Survey
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- RynnEC: Bringing MLLMs into Embodied World
- Agentic Design Review System
- Vision Generalist Model: A Survey
- KFFocus: Highlighting Keyframes for Enhanced Video Understanding
- VGGSounder: Audio-Visual Evaluations for Foundation Models
- TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding
- TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding
- Aligning Effective Tokens with Video Anomaly in Large Language Models
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- EgoTrigger: Toward Audio-Driven Image Capture for Human Memory Enhancement in All-Day Energy-Efficient Smart Glasses
- AffectGPT-R1: Leveraging Reinforcement Learning for Open-Vocabulary Multimodal Emotion Recognition
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
- Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- Gems: Group Emotion Profiling Through Multimodal Situational Understanding
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
- The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- VAGU & GtS: LLM-Based Benchmark and Framework for Joint Video Anomaly Grounding and Understanding
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
- Object-centric Video Question Answering with Visual Grounding and Referring
- Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding
- VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation
- Team of One: Cracking Complex Video QA with Model Synergy
- SV3.3B: A Sports Video Understanding Model for Action Recognition
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- ViTCoT: Video-Text Interleaved Chain-of-Thought for Boosting Video Understanding in Large Language Models
- Smart Routing for Multimodal Video Retrieval: When to Search What
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
- Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- M3-Med: A Benchmark for Multi-lingual, Multi-modal, and Multi-hop Reasoning in Medical Instructional Video Understanding
- EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
- Anti-Prompt: Image Protection against Text-Guided Image-to-Video Generation
- Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
- SciVid: Cross-Domain Evaluation of Video Models in Scientific Applications
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Bisecle: Binding and Separation in Continual Learning for Video Language Understanding
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- EgoM2P: Egocentric Multimodal Multitask Pretraining
- Can Video Large Multimodal Models Think Like Doubters-or Double-Down: A Study on Defeasible Video Entailment
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
- DIVE: Deep-search Iterative Video Exploration A Technical Report for the CVRR Challenge at CVPR 2025
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
- Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
- Fine-Tuning and Prompt Engineering of LLMs, for the Creation of Multi-Agent AI for Addressing Sustainable Protein Production Challenges
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- ToSA: Token Merging with Spatial Awareness
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
- SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
- Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- Speech Recognition on TV Series with Video-guided Post-ASR Correction
- SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- SmartHome-Bench: A Comprehensive Benchmark for Video Anomaly Detection in Smart Homes Using Multi-Modal Large Language Models
- DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
- How Important are Videos for Training Video LLMs?
- CogStream: Context-guided Streaming Video Question Answering
- Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
- Technical Report for Egocentric Mistake Detection for the HoloAssist Challenge
- Movie Facts and Fibs (MF2): A Benchmark for Long Movie Understanding
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
- Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
- Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- Time Blindness: Why Video-Language Models Can't See What Humans Can?
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
- Vid-SME: Membership Inference Attacks against Large Video Understanding Models
- VidText: Towards Comprehensive Evaluation for Video Text Understanding
- VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
- Universal Visuo-Tactile Video Understanding for Embodied Interaction
- Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- HuMoCon: Concept Discovery for Human Motion Understanding
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Reasoning Segmentation for Images and Videos: A Survey
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- ViQAgent: Zero-Shot Video Question Answering via Agent with Open-Vocabulary Grounding Validation
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Domain Adaptation of VLM for Soccer Video Understanding
- Beyond Words: Multimodal LLM Knows When to Speak
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
- A Challenge to Build Neuro-Symbolic Video Agents
- Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
- SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
- LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Temporally-Grounded Language Generation: A Benchmark for Real-Time Vision-Language Models
- Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot
- Generative Pre-trained Autoregressive Diffusion Transformer
- An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
- SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
- Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
- TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
- AdCare-VLM: Towards a Unified and Pre-aligned Latent Representation for Healthcare Video Understanding
- FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
- FSBench: A Figure Skating Benchmark for Advancing Artistic Sports Understanding
- VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
- Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
- ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
- VideoMultiAgents: A Multi-Agent Framework for Video Question Answering
- VEU-Bench: Towards Comprehensive Understanding of Video Editing
- TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
- SurgNarrator: A Generative Retrieval Framework for Surgical Video Understanding
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- Ask2Loc: Learning to Locate Instructional Visual Answers by Asking Questions
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
- ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- RadarLLM: Empowering Large Language Models to Understand Human Motion from Millimeter-Wave Point Cloud Sequence
- AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding
- A Survey on Efficient Vision-Language Models
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
- How Can Objects Help Video-Language Understanding?
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
- PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
- REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
- REVEAL: Relation-based Video Representation Learning for Video-Question-Answering
- Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
- SmolVLM: Redefining small and efficient multimodal models
Discussions
Related