Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
2024/09/18 by Peng Wang, Shuai Bai, Wang, Peng +35 · 1167 citations
Psychology · #Categorization, perception, and language
paper · pdf · doi:10.48550/arxiv.2409.12191
Abstract
We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL introduces the Naive Dynamic Resolution mechanism, which enables the model to dynamically process images of varying resolutions into different numbers of visual tokens. This approach allows the model to generate more efficient and accurate visual representations, closely aligning with human perceptual processes. The model also integrates Multimodal Rotary Position Embedding (M-RoPE), facilitating the effective fusion of positional information across text, images, and videos. We employ a unified paradigm for processing both images and videos, enhancing the model's visual perception capabilities. To explore the potential of large multimodal models, Qwen2-VL investigates the scaling laws for large vision-language models (LVLMs). By scaling both the model size-with versions at 2B, 8B, and 72B parameters-and the amount of training data, the Qwen2-VL Series achieves highly competitive performance. Notably, the Qwen2-VL-72B model achieves results comparable to leading models such as GPT-4o and Claude3.5-Sonnet across various multimodal benchmarks, outperforming other generalist models. Code is available at https://github.com/QwenLM/Qwen2-VL .
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
- RemiAssist: A Therapist-Supporting System for Photo-Based Reminiscence Therapy in Dementia Care
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- LongCat-Image Technical Report
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Active Perception Agent for Omnimodal Audio-Video Understanding
- Instruction-Following Evaluation of Large Vision-Language Models
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
- Video Understanding: From Geometry and Semantics to Unified Models
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- DreamOmni3: Scribble-based Editing and Generation
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
- FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
- MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- STEER: Steerable Dyadic Head Avatars
- Disentangling Semantic Attention from Structural Bias in the Attention Manifold
- Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
- Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
- The Spatial Blindspot of Vision-Language Models
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
- Visual information extraction from documents via classification-guided large vision-language models
- Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- StAR: Segment Anything Reasoner
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Towards Long-window Anchoring in Vision-Language Model Distillation
- Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
- Streaming Video Instruction Tuning
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images
- Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
- Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- Scaling Laws for Energy Efficiency of Local LLMs
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
- The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
- Adapting MLLMs for Nuanced Video Retrieval
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- Motus: A Unified Latent Action World Model
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Using GUI Agent for Electronic Design Automation
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
- FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- VisualActBench: Can VLMs See and Act like a Human?
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- Rethinking Chain-of-Thought Reasoning for Videos
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- Towards Lossless Ultimate Vision Token Compression for VLMs
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- Unified Video Editing with Temporal Reasoner
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- The Role of Entropy in Visual Grounding: Analysis and Optimization
- Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Personalized Image Descriptions from Attention Sequences
- RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- What Happens When: Learning Temporal Orders of Events in Videos
- Concept-based Explainable Data Mining with VLM for 3D Detection
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Jina-VLM: Small Multilingual Vision Language Model
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
- Lumos: Let there be Language Model System Certification
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
- FiMMIA: scaling semantic perturbation-based membership inference across modalities
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- OmniPerson: Unified Identity-Preserving Pedestrian Generation
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- VACoT: Rethinking Visual Data Augmentation with VLMs
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Charts Are Not Images: On the Challenges of Scientific Chart Editing
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
- Optimizing Multimodal Language Models through Attention-based Interpretability
- REVEAL: Reasoning-enhanced Forensic Evidence Analysis for Explainable AI-generated Image Detection
- HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Advancing Aesthetic Image Generation via Composition Transfer
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Qwen3-VL Technical Report
- Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
- AlignBench: Benchmarking Fine-Grained Image-Text Alignment with Synthetic Image-Caption Pairs
- Object-Centric Vision Token Pruning for Vision Language Models
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- CLASH: A Benchmark for Cross-Modal Contradiction Detection
- ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- EEG-VLM: A Hierarchical Vision-Language Model with Multi-Level Feature Alignment and Visually Enhanced Language-Guided Reasoning for EEG Image-Based Sleep Stage Prediction
- Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Understanding Task Transfer in Vision-Language Models
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
- SineProject: Machine Unlearning for Stable Vision Language Alignment
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
- AVERY: Adaptive VLM Split Computing through Embodied Self-Awareness for Efficient Disaster Response Systems
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention
- Attention Guided Alignment in Efficient Vision-Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
- ToC: Tree-of-Claims Search with Multi-Agent Language Models
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Personalized Reward Modeling for Text-to-Image Generation
- Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
- You Only Forward Once: An Efficient Compositional Judging Paradigm
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- KRAL: Knowledge and Reasoning Augmented Learning for LLM-assisted Clinical Antimicrobial Therapy
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Fairness in Multi-modal Medical Diagnosis with Demonstration Selection
- Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
- UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
- GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
- When to Think and When to Look: Uncertainty-Guided Lookback
- Multimodal Evaluation of Russian-language Architectures
- ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- Insight-A: Attribution-aware for Multimodal Misinformation Detection
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- Privacy Preserving Ordinal-Meta Learning with VLMs for Fine-Grained Fruit Quality Prediction
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Decoupled Action Head: Confining Task Knowledge to Conditioning Layers
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
- Suppressing VLM Hallucinations with Spectral Representation Filtering
- MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
- Optimizing Mixture of Block Attention
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- Human-Corrected Labels Learning: Enhancing Labels Quality via Human Correction of VLMs Discrepancies
- CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
- MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- 3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
- Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
- FaithAct: Faithfulness Planning and Acting in MLLMs
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
- Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection
- Streaming Tensor Program: A streaming abstraction for dynamic parallelism
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
- AUTO-Explorer: Automated Data Collection for GUI Agent
- WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
- Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
- S2LM: Towards Semantic Steganography via Large Language Models
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
- A benchmark multimodal oro-dental dataset for large vision-language models
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- Embodiment Transfer Learning for Vision-Language-Action Models
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
- Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
- 3EED: Ground Everything Everywhere in 3D
- Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Reimagining Safety Alignment with An Image
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- Synergistic Tensor and Pipeline Parallelism
- Generating Accurate and Detailed Captions for High-Resolution Images
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- Emu3.5: Native Multimodal Models are World Learners
- Context Engineering 2.0: The Context of Context Engineering
- Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
- FullPart: Generating each 3D Part at Full Resolution
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- More than a Moment: Towards Coherent Sequences of Audio Descriptions
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
- Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
- InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- Vision-Language Models Suppress Female Representations Under Ambiguous Input
- POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- Uniform Discrete Diffusion with Metric Path for Video Generation
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
- VC4VG: Optimizing Video Captions for Text-to-Video Generation
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- Latent Chain-of-Thought for Visual Reasoning
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
- Positional Preservation Embedding for Multimodal Large Language Models
- Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- Agentsway -- Software Development Methodology for AI Agents-based Teams
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- LightAgent: Mobile Agentic Foundation Models
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- [De|Re]constructing VLMs' Reasoning in Counting
- CARES: Context-Aware Resolution Selector for VLMs
- MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression
- olmOCR 2: Unit Test Rewards for Document OCR
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
- The Impact of Image Resolution on Biomedical Multimodal Large Language Models
- Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- HouseTour: A Virtual Real Estate A(I)gent
- Token-Level Inference-Time Alignment for Vision-Language Models
- iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
- From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
- GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Graph4MM: Weaving Multimodal Learning with Structural Information
- An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Directional Reasoning Injection for Fine-Tuning MLLMs
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- RealDPO: Real or Not Real, that is the Preference
- Benchmarking Multimodal Large Language Models for Face Recognition
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- xLLM Technical Report
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- Multimodal Function Vectors for Visual Relations
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Document Intelligence in the Era of Large Language Models: A Survey
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Detect Anything via Next Point Prediction
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- LayerSync: Self-aligning Intermediate Layers
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
- SafeMT: Multi-turn Safety for Multimodal Language Models
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
- video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory
- GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
- Instruction-aware User Embedding via Synergistic Language and Representation Modeling
- Topological Alignment of Shared Vision-Language Embedding Space
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- Task-Aware Resolution Optimization for Visual Large Language Models
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- Auto-scaling Continuous Memory for GUI Agent
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
- LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
- Leveraging LLMs to Streamline the Review of Public Funding Applications
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
- ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- DreamOmni2: Multimodal Instruction-based Editing and Generation
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- The Safety Challenge of World Models for Embodied AI Agents: A Review
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Visual Representations inside the Language Model
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
- Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- IMAGEdit: Let Any Subject Transform
- ModernVBERT: Towards Smaller Visual Document Retrievers
- What You See is What You Ask: Evaluating Audio Descriptions
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction
- OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- DGM4+: Dataset Extension for Global Scene Inconsistency
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- Vision Function Layer in Multimodal LLMs
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UI-UG: A Unified MLLM for UI Understanding and Generation
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
- UniVid: The Open-Source Unified Video Model
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- NeMo: Needle in a Montage for Video-Language Understanding
- FreeRet: MLLMs as Training-Free Retrievers
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
- SVAC: Scaling Is All You Need For Referring Video Object Segmentation
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
- Video Panels for Long Video Understanding
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
- MMPB: It's Time for Multi-Modal Personalization
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
- Explaining multimodal LLMs via intra-modal token interactions
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- Multilingual Vision-Language Models, A Survey
- FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- Revisiting Data Challenges of Computational Pathology: A Pack-based Multiple Instance Learning Training Framework
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
- Confidence-guided Refinement Reasoning for Zero-shot Question Answering
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Logics-Parsing Technical Report
- Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
- Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
- Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
- Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
- Gaze Heads: How VLMs Look at What They Describe
- S-GRPO: Unified Post-Training for Large Vision-Language Models
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- Use What You Know: Causal Foundation Models with Partial Graphs
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
- F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
- Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
- Mano Technical Report
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- Memory-QA: Answering Recall Questions Based on Multimodal Memories
- ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
- Modeling Bottom-up Information Quality during Language Processing
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Can GRPO Boost Complex Multimodal Table Understanding?
- A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- AToken: A Unified Tokenizer for Vision
- Dense Video Understanding with Gated Residual Tokenization
- Lightweight Joint Optimization of General-Purpose Vision-Language Models and Retrievers for RAG-Based Medical Diagnosis
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
- CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Chat-Driven Text Generation and Interaction for Person Retrieval
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- 3D Aware Region Prompted Vision Language Model
- Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
- Improving Fungi Prototype Representations for Few-Shot Classification
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
- Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
- LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
- Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Visual Grounding from Event Cameras
- Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- AdsQA: Towards Advertisement Video Understanding
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
- No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
- Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
- TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- KLIPA: A Knowledge Graph and LLM-Driven QA Framework for IP Analysis
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Semantic-guided LoRA Parameters Generation
- ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
- Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
- Planning with Reasoning using Vision Language World Model
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
- Reinforced Visual Perception with Tools
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Delta Rectified Flow Sampling for Text-to-Image Editing
- Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models
- Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
- Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- MedGR2: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- SoccerNet 2025 Challenges Results
- Rethinking Human-Object Interaction Evaluation for both Vision-Language Models and HOI-Specific Methods
- PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
- Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
- DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- Region-Level Context-Aware Multimodal Understanding
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
- OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Language models align with brain regions that represent concepts across modalities
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- Towards Agentic AI for Multimodal-Guided Video Object Segmentation
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- Preacher: Paper-to-Video Agentic System
- SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
- Episodic Memory Representation for Long-form Video Understanding
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- OpenCUA: Open Foundations for Computer-Use Agents
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
- CARES: Collaborative Agentic Reasoning for Error Detection in Surgery
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility
- RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
- Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Δ-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
- Text-guided Visual Prompt DINO for Generic Segmentation
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- DreamVE: Unified Instruction-based Image and Video Editing
- Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
- IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection
- QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- Training-Free Multimodal Large Language Model Orchestration
- Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
- HPSv3: Towards Wide-Spectrum Human Preference Score
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
- Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- PixNerd: Pixel Neural Field Diffusion
- Adversarial-Guided Diffusion for Multimodal LLM Attacks
- On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
- BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
- Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
- PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems
- SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
- Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
- Object-centric Video Question Answering with Visual Grounding and Referring
- A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
- A Survey of Multimodal Hallucination Evaluation and Detection
- LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
- True Multimodal In-Context Learning Needs Attention to the Visual Context
- Advancing Complex Video Object Segmentation via Progressive Concept Construction
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
- VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
- Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression
- Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark
- InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
- VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
- Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
- VLA-Mark: A cross modal watermark for large vision-language alignment model
- PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
- Eyes Will Shut: A Vision-Based Next GPS Location Prediction Model by Reinforcement Learning from Visual Map Feed Back
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
- Mitigating Object Hallucinations via Sentence-Level Early Intervention
- Human-like object concept representations emerge naturally in multimodal large language models
- POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
- NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
- Demonstrating the Octopi-1.5 Visual-Tactile-Language Model
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
- VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
- InsurTech innovation using natural language processing
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- Learning Diffusion Models with Flexible Representation Guidance
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Multilingual Multimodal Software Developer for Code Generation
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
- A document is worth a structured record: Principled inductive bias design for document recognition
- Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
- VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models
- BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
- MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- Scaling RL to Long Videos
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models
- Omni-Video: Democratizing Unified Video Understanding and Generation
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation
- Skywork-R1V3 Technical Report
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Transcribing Spanish Texts from the Past: Experiments with Transkribus, Tesseract and Granite
- From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
- Training-free Generation of Temporally Consistent Rewards from VLMs
- From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
- Vision-Language Models Can't See the Obvious
- Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
- Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
- PresentAgent: Multimodal Agent for Presentation Video Generation
- LLM4Hint: Leveraging Large Language Models for Hint Recommendation in Offline Query Optimization
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- No time to train! Training-Free Reference-Based Instance Segmentation
- From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
- LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
- Activation Reward Models for Few-Shot Model Alignment
- AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
- Just Noticeable Difference for Large Multimodal Models
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- Token Activation Map to Visually Explain Multimodal LLMs
- VisualPrompter: Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
- Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
- UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
- Ovis-U1 Technical Report
- ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
- Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- Test-Time Consistency in Vision Language Models
- Rethinking Visual Token Reduction in LVLMs under Cross-modal Misalignment
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
- Universal Retrieval for Multimodal Trajectory Modeling
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
- MDC-R: The Minecraft Dialogue Corpus with Reference
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- Class-Agnostic Region-of-Interest Matching in Document Images
- LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
- Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards
- Visual-Semantic Knowledge Conflicts in Operating Rooms: Synthetic Data Curation for Surgical Risk Perception in Multimodal Large Language Models
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
- Unified Vision-Language-Action Model
- MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
- Capturing Fine-Grained Alignments Improves 3D Affordance Detection
- Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
- Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
- OmniGen2: Exploration to Advanced Multimodal Generation
- Object-aware Sound Source Localization via Audio-Visual Scene Understanding
- Generalizing vision-language models to novel domains: A comprehensive survey
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
- Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
- MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
- CLGRPO: Reasoning Ability Enhancement for Small VLMs
- Adapting Vision-Language Models for Evaluating World Models
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
- PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
- PhysUniBench: An Undergraduate-Level Physics Reasoning Benchmark for Multimodal Models
- MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- Structured Attention Matters to Multimodal LLMs in Document Understanding
- Visual symbolic mechanisms: Emergent symbol processing in vision language models
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
Related