Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
2024/09/18 by Peng Wang, Shuai Bai, Wang, Peng +35 · 1,828 citations
Psychology · #Categorization, perception, and language
paper · pdf · doi:10.48550/arxiv.2409.12191
Abstract
We present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing. Qwen2-VL introduces the Naive Dynamic Resolution mechanism, which enables the model to dynamically process images of varying resolutions into different numbers of visual tokens. This approach allows the model to generate more efficient and accurate visual representations, closely aligning with human perceptual processes. The model also integrates Multimodal Rotary Position Embedding (M-RoPE), facilitating the effective fusion of positional information across text, images, and videos. We employ a unified paradigm for processing both images and videos, enhancing the model's visual perception capabilities. To explore the potential of large multimodal models, Qwen2-VL investigates the scaling laws for large vision-language models (LVLMs). By scaling both the model size-with versions at 2B, 8B, and 72B parameters-and the amount of training data, the Qwen2-VL Series achieves highly competitive performance. Notably, the Qwen2-VL-72B model achieves results comparable to leading models such as GPT-4o and Claude3.5-Sonnet across various multimodal benchmarks, outperforming other generalist models. Code is available at https://github.com/QwenLM/Qwen2-VL .
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
- RemiAssist: A Therapist-Supporting System for Photo-Based Reminiscence Therapy in Dementia Care
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- LongCat-Image Technical Report
- Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
- Active Perception Agent for Omnimodal Audio-Video Understanding
- Instruction-Following Evaluation of Large Vision-Language Models
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- TV-RAG: A Temporal-aware and Semantic Entropy-Weighted Framework for Long Video Retrieval and Understanding
- Video Understanding: From Geometry and Semantics to Unified Models
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- DreamOmni3: Scribble-based Editing and Generation
- MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video Generation
- FETAL-GAUGE: A Benchmark for Assessing Vision-Language Models in Fetal Ultrasound
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- LeapBot-WA: World-Anchor Action Models via Predictive Latent Alignments
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
- MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- Being-H0.7: A Latent World-Action Model from Egocentric Videos
- STEER: Steerable Dyadic Head Avatars
- Disentangling Semantic Attention from Structural Bias in the Attention Manifold
- Are Prompt Optimizers Blind? Cross-Modal Visual Feedback for Automatic Prompt Optimization
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation
- Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
- The Spatial Blindspot of Vision-Language Models
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- PCA: Persistence-Aware Compression and Aggregation for Fast Video Large Language Models
- Visual information extraction from documents via classification-guided large vision-language models
- Child-Oriented AIGC Video Risk Reviewing: A Benchmark and Knowledge-Supported Iterative Reasoning Framework
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- StAR: Segment Anything Reasoner
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Towards Long-window Anchoring in Vision-Language Model Distillation
- Towards Responsible and Explainable AI Agents with Consensus-Driven Reasoning
- Streaming Video Instruction Tuning
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images
- Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
- LongVideoAgent: Multi-Agent Reasoning with Long Videos
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Widget2Code: From Visual Widgets to UI Code via Multimodal LLMs
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- Generative Giants, Retrieval Weaklings: Why do Multimodal Large Language Models Fail at Multimodal Retrieval?
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- Restore-R1: Efficient Image Restoration Agents via Reinforcement Learning with Multimodal LLM Perceptual Feedback
- Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation
- Layout-Aware Text Editing for Efficient Transformation of Academic PDFs to Markdown
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- Exploration of Augmentation Strategies in Multi-modal Retrieval-Augmented Generation for the Biomedical Domain: A Case Study Evaluating Question Answering in Glycobiology
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- Scaling Laws for Energy Efficiency of Local LLMs
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Seeing is Believing (and Predicting): Context-Aware Multi-Human Behavior Prediction with Vision Language Models
- The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- ViRC: Enhancing Visual Interleaved Mathematical CoT with Reason Chunking
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
- SketchAssist: A Practical Assistant for Semantic Edits and Precise Local Redrawing
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
- HyperVL: An Efficient and Dynamic Multimodal Large Language Model for Edge Devices
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- EchoVLM: Measurement-Grounded Multimodal Learning for Echocardiography
- VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
- Adapting MLLMs for Nuanced Video Retrieval
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- Motus: A Unified Latent Action World Model
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- Using GUI Agent for Electronic Design Automation
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Limits and Gains of Test-Time Scaling in Vision-Language Reasoning
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Empowering Dynamic Urban Navigation with Stereo and Mid-Level Vision
- FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- ShotDirector: Directorially Controllable Multi-Shot Video Generation with Cinematographic Transitions
- VisualActBench: Can VLMs See and Act like a Human?
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
- Rethinking Chain-of-Thought Reasoning for Videos
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- View-on-Graph: Zero-shot 3D Visual Grounding via Vision-Language Reasoning on Scene Graphs
- DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
- A Practical Guide for Designing, Developing, and Deploying Production-Grade Agentic AI Workflows
- Towards Lossless Ultimate Vision Token Compression for VLMs
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- Beyond Real: Imaginary Extension of Rotary Position Embeddings for Long-Context LLMs
- Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- VideoCoF: Unified Video Editing with Temporal Reasoner
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- The Role of Entropy in Visual Grounding: Analysis and Optimization
- Statistic-Augmented, Decoupled MoE Routing and Aggregating in Autonomous Driving
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Personalized Image Descriptions from Attention Sequences
- RunawayEvil: Jailbreaking the Image-to-Video Generative Models
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- Are AI-Generated Driving Videos Ready for Autonomous Driving? A Diagnostic Evaluation Framework
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- Zoom in, Click out: Unlocking and Evaluating the Potential of Zooming for GUI Grounding
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- What Happens When: Learning Temporal Orders of Events in Videos
- Concept-based Explainable Data Mining with VLM for 3D Detection
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- Semore: VLM-guided Enhanced Semantic Motion Representations for Visual Reinforcement Learning
- GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
- Jina-VLM: Small Multilingual Vision Language Model
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- M3DR: Towards Universal Multilingual Multimodal Document Retrieval
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- Fairness-Aware Fine-Tuning of Vision-Language Models for Medical Glaucoma Diagnosis
- PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
- Lumos: Let there be Language Model System Certification
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Making Dialogue Grounding Data Rich: A Three-Tier Data Synthesis Framework for Generalized Referring Expression Comprehension
- FiMMIA: scaling semantic perturbation-based membership inference across modalities
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- OmniPerson: Unified Identity-Preserving Pedestrian Generation
- WISE: Weighted Iterative Society-of-Experts for Robust Multimodal Multi-Agent Debate
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- VACoT: Rethinking Visual Data Augmentation with VLMs
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models
- ChartAnchor: Chart Grounding with Structural-Semantic Fidelity
- Augmenting Intra-Modal Understanding in MLLMs for Robust Multimodal Keyphrase Generation
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Charts Are Not Images: On the Challenges of Scientific Chart Editing
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- RealGen: Photorealistic Text-to-Image Generation via Detector-Guided Rewards
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Hunyuan-GameCraft-2: Instruction-following Interactive Game World Model
- Optimizing Multimodal Language Models through Attention-based Interpretability
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
- HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- Adapting Like Humans: A Metacognitive Agent with Test-time Reasoning
- UniGeoSeg: Towards Unified Open-World Segmentation for Geospatial Scenes
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- Advancing Aesthetic Image Generation via Composition Transfer
- Canvas-to-Image: Compositional Image Generation with Multimodal Controls
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- PROMPTMINER: Black-Box Prompt Stealing against Text-to-Image Generative Models via Reinforcement Learning and Fuzz Optimization
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
- FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Qwen3-VL Technical Report
- Beyond Generation: Multi-Hop Reasoning for Factual Accuracy in Vision-Language Models
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- Object-Centric Vision Token Pruning for Vision Language Models
- Text-guided Controllable Diffusion for Realistic Camouflage Images Generation
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- CLASH: A Benchmark for Cross-Modal Contradiction Detection
- DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
- ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- EEG-VLM: A Hierarchical Vision-Language Model with Multi-Level Feature Alignment and Visually Enhanced Language-Guided Reasoning for EEG Image-Based Sleep Stage Prediction
- Human-Centric Open-Future Task Discovery: Formulation, Benchmark, and Scalable Tree-Based Search
- EventSTU: Event-Guided Efficient Spatio-Temporal Understanding for Video Large Language Models
- MagicWorld: Towards Long-Horizon Stability for Interactive Video World Exploration
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- Understanding Task Transfer in Vision-Language Models
- Multimodal Large Language Models with Adaptive Preference Optimization for Sequential Recommendation
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
- SineProject: Machine Unlearning for Stable Vision Language Alignment
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- WebSTAR: Scalable Data Synthesis for Computer Use Agents with Step-Level Filtering
- AVERY: Adaptive VLM Split Computing through Embodied Self-Awareness for Efficient Disaster Response Systems
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention
- Attention Guided Alignment in Efficient Vision-Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
- ToC: Tree-of-Claims Search with Multi-Agent Language Models
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Personalized Reward Modeling for Text-to-Image Generation
- Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
- Revisiting Multimodal KV Cache Compression: A Frequency-Domain-Guided Outlier-KV-Aware Approach
- You Only Forward Once: An Efficient Compositional Judging Paradigm
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- KRAL: Knowledge and Reasoning Augmented Learning for LLM-assisted Clinical Antimicrobial Therapy
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Fairness in Multi-modal Medical Diagnosis with Demonstration Selection
- Mind the Motions: Benchmarking Theory-of-Mind in Everyday Body Language
- UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
- GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization
- When to Think and When to Look: Uncertainty-Guided Lookback
- Multimodal Evaluation of Russian-language Architectures
- ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
- Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
- O3SLM: Open Weight, Open Data, and Open Vocabulary Sketch-Language Model
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- Insight-A: Attribution-aware for Multimodal Misinformation Detection
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- Privacy Preserving Ordinal-Meta Learning with VLMs for Fine-Grained Fruit Quality Prediction
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Decoupled Action Expert: Confining Task Knowledge to the Conditioning Pathway
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
- Suppressing VLM Hallucinations with Spectral Representation Filtering
- MAVIS: A Benchmark for Multimodal Source Attribution in Long-form Visual Question Answering
- Adaptive Diagnostic Reasoning Framework for Pathology with Multimodal Large Language Models
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- LIHE: Linguistic Instance-Split Hyperbolic-Euclidean Framework for Generalized Weakly-Supervised Referring Expression Comprehension
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models
- AirCopBench: A Benchmark for Multi-drone Collaborative Embodied Perception and Reasoning
- PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
- Spatial Reasoning in Multimodal Large Language Models: A Survey of Tasks, Benchmarks and Methods
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation
- Optimizing Mixture of Block Attention
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- Human-Corrected Labels Learning: Enhancing Labels Quality via Human Correction of VLMs Discrepancies
- CrochetBench: Can Vision-Language Models Move from Describing to Doing in Crochet Domain?
- MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- 3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
- Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
- Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
- VideoChain: A Transformer-Based Framework for Multi-hop Video Question Generation
- Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
- Multi-Modal Assistance for Unsupervised Domain Adaptation on Point Cloud 3D Object Detection
- Streaming Tensor Program: A streaming abstraction for dynamic parallelism
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- A Circular Argument : Does RoPE need to be Equivariant for Vision?
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning
- Cross-Modal Unlearning via Influential Neuron Path Editing in Multimodal Large Language Models
- Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
- AUTO-Explorer: Automated Data Collection for GUI Agent
- WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
- Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic Interpretability and Hallucination Mitigation
- S2LM: Towards Semantic Steganography via Large Language Models
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
- A benchmark multimodal oro-dental dataset for large vision-language models
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- Embodiment Transfer Learning for Vision-Language-Action Models
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Contamination Detection for VLMs using Multi-Modal Semantic Perturbation
- GUIDES: Guidance Using Instructor-Distilled Embeddings for Pre-trained Robot Policy Enhancement
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- DetectiumFire: A Comprehensive Multi-modal Dataset Bridging Vision and Language for Fire Understanding
- ChartM3: A Multi-Stage Code-Driven Pipeline for Constructing Multi-Dimensional and Multi-Step Visual Reasoning Data in Chart Comprehension
- Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
- 3EED: Ground Everything Everywhere in 3D
- Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Reimagining Safety Alignment with An Image
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Enhancing Trustworthy GUI Grounding via Self-Critiqued Reinforcement Learning
- Synergistic Tensor and Pipeline Parallelism
- Generating Accurate and Detailed Captions for High-Resolution Images
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- BEAT: Visual Backdoor Attacks on VLM-based Embodied Agents via Contrastive Trigger Learning
- GeoFM: Enhancing Geometric Reasoning of MLLMs via Synthetic Data Generation through Formal Language
- Emu3.5: Native Multimodal Models are World Learners
- Context Engineering 2.0: The Context of Context Engineering
- Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- Which Way Does Time Flow? A Psychophysics-Grounded Evaluation for Vision-Language Models
- FullPart: Generating each 3D Part at Full Resolution
- Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- Standardization of Psychiatric Diagnoses -- Role of Fine-tuned LLM Consortium and OpenAI-gpt-oss Reasoning LLM Enabled Decision Support System
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- More than a Moment: Towards Coherent Sequences of Audio Descriptions
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- Anchoring and Steering Diffusion: Enhancing the Faithfulness of Text-to-Image Generation at Inference Time
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
- ChartReasoner: Code-Driven Modality Bridging for Long-Chain Reasoning in Chart Question Answering
- Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
- InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- Vision-Language Models Suppress Female Representations Under Ambiguous Input
- POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- Uniform Discrete Diffusion with Metric Path for Video Generation
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- SPARTA: Evaluating Reasoning Segmentation Robustness through Black-Box Adversarial Paraphrasing in Text Autoencoder Latent Space
- Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- DynaRend: Learning 3D Dynamics via Masked Future Rendering for Robotic Manipulation
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- MuSaG: A Multimodal German Sarcasm Dataset with Full-Modal Annotations
- VC4VG: Optimizing Video Captions for Text-to-Video Generation
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- Latent Chain-of-Thought for Visual Reasoning
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- Emotion-Coherent Reasoning for Multimodal LLMs via Emotional Rationale Verifier
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
- Positional Preservation Embedding for Multimodal Large Language Models
- Rethinking the Text-Vision Reasoning Imbalance in MLLMs through the Lens of Training Recipes
- S-Chain: Structured Visual Chain-of-Thought For Medicine
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- Agentsway -- Software Development Methodology for AI Agents-based Teams
- PromptReverb: Multimodal Room Impulse Response Generation Through Latent Rectified Flow Matching
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- LightAgent: Mobile Agentic Foundation Models
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- A Multimodal Benchmark for Framing of Oil & Gas Advertising and Potential Greenwashing Detection
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Multi-Step Reasoning for Embodied Question Answering via Tool Augmentation
- Why LVLMs Are More Prone to Hallucinations in Longer Responses: The Role of Context
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- [De|Re]constructing VLMs' Reasoning in Counting
- CARES: Context-Aware Resolution Selector for VLMs
- MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression
- olmOCR 2: Unit Test Rewards for Document OCR
- VITA-E: Natural Embodied Interaction with Concurrent Seeing, Hearing, Speaking, and Acting
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- Activating Visual Context and Commonsense Reasoning through Masked Prediction in VLMs
- The Impact of Image Resolution on Biomedical Multimodal Large Language Models
- Text or Pixels? It Takes Half: On the Token Efficiency of Visual Text Inputs in Multimodal LLMs
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- RadDiagSeg-M: A Vision Language Model for Joint Diagnosis and Multi-Target Segmentation in Radiology
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- DeepSeek-OCR: Contexts Optical Compression
- UWBench: A Comprehensive Vision-Language Benchmark for Underwater Understanding
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- HouseTour: A Virtual Real Estate A(I)gent
- Token-Level Inference-Time Alignment for Vision-Language Models
- iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
- From Preferences to Prejudice: The Role of Alignment Tuning in Shaping Social Bias in Video Diffusion Models
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- Enhanced Motion Forecasting with Plug-and-Play Multimodal Large Language Models
- VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs
- GACO-CAD: Geometry-Augmented and Conciseness-Optimized CAD Model Generation from Single Image
- SARSteer: Safeguarding Large Audio-Language Models via Safe-Ablated Refusal Steering
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Graph4MM: Weaving Multimodal Learning with Structural Information
- An Efficient Framework for Whole-Page Reranking via Single-Modal Supervision
- Does Visual Grounding Enhance the Understanding of Embodied Knowledge in Large Language Models?
- VERITAS: Leveraging Vision Priors and Expert Fusion to Improve Multimodal Data
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- Directional Reasoning Injection for Fine-Tuning MLLMs
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- RealDPO: Real or Not Real, that is the Preference
- Benchmarking Multimodal Large Language Models for Face Recognition
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- xLLM Technical Report
- VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
- Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- MorphoBench: A Benchmark with Difficulty Adaptive to Model Reasoning
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- Multimodal Function Vectors for Visual Relations
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Document Intelligence in the Era of Large Language Models: A Survey
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
- UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity MoE
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Detect Anything via Next Point Prediction
- Reasoning in the Dark: Interleaved Vision-Text Reasoning in Latent Space
- LayerSync: Self-aligning Intermediate Layers
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- Hierarchical Reasoning with Vision-Language Models for Incident Reports from Dashcam Videos
- SafeMT: Multi-turn Safety for Multimodal Language Models
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- mmWalk: Towards Multi-modal Multi-view Walking Assistance
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web Coding
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
- video-SALMONN S: Streaming Audio-Visual LLMs Beyond Length Limits via Memory
- GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
- Instruction-aware User Embedding via Synergistic Language and Representation Modeling
- Topological Alignment of Shared Vision-Language Embedding Space
- Where on Earth? A Vision-Language Benchmark for Probing Model Geolocation Skills Across Scales
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Unified Open-World Segmentation with Multi-Modal Prompts
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- Task-Aware Resolution Optimization for Visual Large Language Models
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- Auto-scaling Continuous Memory for GUI Agent
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- Speculative Jacobi-Denoising Decoding for Accelerating Autoregressive Text-to-image Generation
- ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling
- LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- Beyond Over-Refusal: Scenario-Based Diagnostics and Post-Hoc Mitigation for Exaggerated Refusals in LLMs
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- D-CoDe: Scaling Image-Pretrained VLMs to Video via Dynamic Compression and Question Decomposition
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- Deploying Tiny LVLM Judges for Real-World Evaluation of Chart Models: Lessons Learned and Best Practices
- Leveraging LLMs to Streamline the Review of Public Funding Applications
- Are We Using the Right Benchmark: An Evaluation Framework for Visual Token Compression Methods
- ExPO-HM: Learning to Explain-then-Detect for Hateful Meme Detection
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- FailSafe: Reasoning and Recovery from Failures in Vision-Language-Action Models
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- DreamOmni2: Multimodal Instruction-based Editing and Generation
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- VideoMiner: Iteratively Grounding Key Frames of Hour-Long Videos via Tree-based Group Relative Policy Optimization
- The Safety Challenge of World Models for Embodied AI Agents: A Review
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- AgeBooth: Controllable Facial Aging and Rejuvenation via Diffusion Models
- OASIS: A Multilingual and Multimodal Dataset for Culturally Grounded Spoken Visual QA
- FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
- Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Visual Representations inside the Language Model
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- Mitigating Diffusion Model Hallucinations with Dynamic Guidance
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
- Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- No Tokens Wasted: Leveraging Long Context in Biomedical Vision-Language Models
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- AdaRD-key: Adaptive Relevance-Diversity Keyframe Sampling for Long-form Video understanding
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- IMAGEdit: Let Any Subject Transform
- ModernVBERT: Towards Smaller Visual Document Retrievers
- What You See is What You Ask: Evaluating Audio Descriptions
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction
- OTTER: Open-Tagging via Text-Image Representation for Multi-modal Understanding
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- DGM4+: Dataset Extension for Global Scene Inconsistency
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- V-HUB: A Visual-Centric Humor Understanding Benchmark for Video LLMs
- NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language
- Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- A Navigation Framework Utilizing Vision-Language Models
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- Vision Function Layer in Multimodal LLMs
- Can you SPLICE it together? A Human Curated Benchmark for Probing Visual Reasoning in VLMs
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- UI-UG: A Unified MLLM for UI Understanding and Generation
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Bridging the behavior-neural gap: A multimodal AI reveals the brain's geometry of emotion more accurately than human self-reports
- UniVid: The Open-Source Unified Video Model
- Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding
- NeMo: Needle in a Montage for Video-Language Understanding
- FreeRet: MLLMs as Training-Free Retrievers
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
- SVAC: Scaling Is All You Need For Referring Video Object Segmentation
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
- Video Panels for Long Video Understanding
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- WorldSplat: Gaussian-Centric Feed-Forward 4D Scene Generation for Autonomous Driving
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
- Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
- MMPB: It's Time for Multi-Modal Personalization
- VLA-Reasoner: Empowering Vision-Language-Action Models with Reasoning via Online Monte Carlo Tree Search
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
- Explaining multimodal LLMs via intra-modal token interactions
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- UniMapGen: A Generative Framework for Large-Scale Map Construction from Multi-modal Data
- MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
- Multilingual Vision-Language Models, A Survey
- FailureAtlas:Mapping the Failure Landscape of T2I Models via Active Exploration
- ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- MIRG-RL: Multi-Image Reasoning and Grounding with Reinforcement Learning
- UISim: An Interactive Image-Based UI Simulator for Dynamic Mobile Environments
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- X-CoT: Explainable Text-to-Video Retrieval via LLM-based Chain-of-Thought Reasoning
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- MOSS-ChatV: Reinforcement Learning with Process Reasoning Reward for Video Temporal Reasoning
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- Revisiting Data Challenges of Computational Pathology: A Pack-based Multiple Instance Learning Training Framework
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- Recon-Act: A Self-Evolving Multi-Agent Browser-Use System via Web Reconnaissance, Tool Generation, and Task Execution
- Confidence-guided Refinement Reasoning for Zero-shot Question Answering
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Logics-Parsing Technical Report
- Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
- Anatomy of a Feeling: Narrating Embodied Emotions via Large Vision-Language Models
- Calibrate Before Reason: Robust Visual Token Reduction against Semantic Drift in VLMs
- Look But Don't Touch with Sparse Autoencoders for Unlearning in Diffusion Models
- Gaze Heads: How VLMs Look at What They Describe
- S-GRPO: Unified Post-Training for Large Vision-Language Models
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- Use What You Know: Causal Foundation Models with Partial Graphs
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- VISA: Group-wise Visual Token Selection and Aggregation via Graph Summarization for Efficient MLLMs Inference
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- DRISHTIKON: A Multimodal Multilingual Benchmark for Testing Language Models' Understanding on Indian Culture
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained Rewards
- Hyper-Bagel: A Unified Acceleration Framework for Multimodal Understanding and Generation
- RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
- F2RVLM: Boosting Fine-grained Fragment Retrieval for Multi-Modal Long-form Dialogue with Vision Language Model
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- Live-E2T: Real-time Threat Monitoring in Video via Deduplicated Event Reasoning and Chain-of-Thought
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
- Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- RealBench: A Chinese Multi-image Understanding Benchmark Close to Real-world Scenarios
- Mano Technical Report
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- MLLM-Driven Semantic Identifier Generation for Generative Cross-Modal Retrieval
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- Memory-QA: Answering Recall Questions Based on Multimodal Memories
- ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
- Modeling Bottom-up Information Quality during Language Processing
- From Easy to Hard: The MIR Benchmark for Progressive Interleaved Multi-Image Reasoning
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Can GRPO Boost Complex Multimodal Table Understanding?
- A Chain-of-thought Reasoning Breast Ultrasound Dataset Covering All Histopathology Categories
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- Learning from Gene Names, Expression Values and Images: Contrastive Masked Text-Image Pretraining for Spatial Transcriptomics Representation Learning
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- When Big Models Train Small Ones: Label-Free Model Parity Alignment for Efficient Visual Question Answering using Small VLMs
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
- ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Pointing to a Llama and Call it a Camel: On the Sycophancy of Multimodal Large Language Models
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Beyond Spurious Signals: Debiasing Multimodal Large Language Models via Counterfactual Inference and Adaptive Expert Routing
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- Ask-to-Clarify: Resolving Instruction Ambiguity through Multi-turn Dialogue
- EchoVLM: Dynamic Mixture-of-Experts Vision-Language Model for Universal Ultrasound Intelligence
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- DF-LLaVA: Unlocking MLLMs for Synthetic Image Detection via Knowledge Injection and Conflict-Driven Self-Reflection
- AToken: A Unified Tokenizer for Vision
- Dense Video Understanding with Gated Residual Tokenization
- Lightweight Joint Optimization of General-Purpose Vision-Language Models and Retrievers for RAG-Based Medical Diagnosis
- SSL-SSAW: Self-Supervised Learning with Sigmoid Self-Attention Weighting for Question-Based Sign Language Translation
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
- CultranAI at PalmX 2025: Data Augmentation for Cultural Knowledge Representation
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Chat-Driven Text Generation and Interaction for Person Retrieval
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- 3D Aware Region Prompted Vision Language Model
- Small Models, Big Results: Achieving Superior Intent Extraction through Decomposition
- OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- Mitigating Hallucinations in Large Vision-Language Models by Self-Injecting Hallucinations
- Improving Fungi Prototype Representations for Few-Shot Classification
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- OpenHA: A Series of Open-Source Hierarchical Agentic Models in Minecraft
- GAMMA: Generalizable Alignment via Multi-task and Manipulation-Augmented Training for AI-Generated Image Detection
- Robust Diagram Reasoning: A Framework for Enhancing LVLM Performance on Visually Perturbed Scientific Diagrams
- LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
- Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Visual Grounding from Event Cameras
- Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- AdsQA: Towards Advertisement Video Understanding
- Beyond Emotion Recognition: A Multi-Turn Multimodal Emotion Understanding and Reasoning Benchmark
- MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
- No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
- Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
- TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- KLIPA: A Knowledge Graph and LLM-Driven QA Framework for IP Analysis
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- MoLoRAG: Bootstrapping Document Understanding via Multi-modal Logic-aware Retrieval
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Semantic-guided LoRA Parameters Generation
- ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
- DreamPRM-1.5: Unlocking the Potential of Each Instance for Multimodal Process Reward Model Training
- TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
- Self-adaptive Dataset Construction for Real-World Multimodal Safety Scenarios
- Promptception: How Sensitive Are Large Multimodal Models to Prompts?
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- InstaDA: Augmenting Instance Segmentation Data with Dual-Agent System
- Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
- Planning with Reasoning using Vision Language World Model
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- RSCC: A Large-Scale Remote Sensing Change Caption Dataset for Disaster Events
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- VaccineRAG: Boosting Multimodal Large Language Models' Immunity to Harmful RAG Samples
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Succeed or Learn Slowly: Sample Efficient Off-Policy Reinforcement Learning for Mobile App Control
- Reinforced Visual Perception with Tools
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- POINTS-Reader: Distillation-Free Adaptation of Vision-Language Models for Document Conversion
- FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Delta Rectified Flow Sampling for Text-to-Image Editing
- Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models
- Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
- LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- Image-to-Brain Signal Generation for Visual Prosthesis with CLIP Guided Multimodal Diffusion Models
- One VLM, Two Roles: Stage-Wise Routing and Specialty-Level Deployment for Clinical Workflows
- Activation Steering Meets Preference Optimization: Defense Against Jailbreaks in Vision Language Models
- TrimTokenator: Towards Adaptive Visual Token Pruning for Large Multimodal Models
- VideoRewardBench: Comprehensive Evaluation of Multimodal Reward Models for Video Understanding
- LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
- KG-RAG: Enhancing GUI Agent Decision-Making via Knowledge Graph-Driven Retrieval-Augmented Generation
- Safe-LLaVA: A Privacy-Preserving Vision-Language Dataset and Benchmark for Biometric Safety
- TMUAD: Enhancing Logical Capabilities in Unified Anomaly Detection Models with a Text Memory Bank
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
- Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- MedGR2: Breaking the Data Barrier for Medical Reasoning via Generative Reward Learning
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- NLKI: A lightweight Natural Language Knowledge Integration Framework for Improving Small VLMs in Commonsense VQA Tasks
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- SoccerNet 2025 Challenges Results
- CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
- PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
- Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- A Language-Signal-Vision Multimodal Framework for Multitask Cardiac Analysis
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- ViLaD: A Large Vision Language Diffusion Framework for End-to-End Autonomous Driving
- DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- EGOILLUSION: Benchmarking Hallucinations in Egocentric Video Understanding
- ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- Region-Level Context-Aware Multimodal Understanding
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- Standardization of Neuromuscular Reflex Analysis -- Role of Fine-Tuned Vision-Language Model Consortium and OpenAI gpt-oss Reasoning LLM Enabled Decision Support System
- VimoRAG: Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- ExploreVLM: Closed-Loop Robot Exploration Task Planning with Vision-Language Models
- OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Language models align with brain regions that represent concepts across modalities
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- UI-Venus Technical Report: Building High-performance UI Agents with RFT
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- IADGPT: Unified LVLM for Few-Shot Industrial Anomaly Detection, Localization, and Reasoning via In-Context Learning
- HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
- Towards Agentic AI for Multimodal-Guided Video Object Segmentation
- Vision Generalist Model: A Survey
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- Pruning Long Chain-of-Thought of Large Reasoning Models via Small-Scale Preference Optimization
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement
- Preacher: Paper-to-Video Agentic System
- SOI is the Root of All Evil: Quantifying and Breaking Similar Object Interference in Single Object Tracking
- Episodic Memory Representation for Long-form Video Understanding
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- HumanGenesis: Agent-Based Geometric and Generative Modeling for Synthetic Human Dynamics
- RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- OpenCUA: Open Foundations for Computer-Use Agents
- MVISU-Bench: Benchmarking Mobile Agents for Real-World Tasks by Multi-App, Vague, Interactive, Single-App and Unethical Instructions
- CARES: Collaborative Agentic Reasoning for Error Detection in Surgery
- STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision
- The Escalator Problem: Identifying Implicit Motion Blindness in AI for Accessibility
- RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering
- Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- CATP: Contextually Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning
- Small-Large Collaboration: Training-efficient Concept Personalization for Large VLM using a Meta Personalized Small VLM
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Δ-AttnMask: Attention-Guided Masked Hidden States for Efficient Data Selection and Augmentation
- Text-guided Visual Prompt DINO for Generic Segmentation
- SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
- AdaptInfer: Adaptive Token Pruning for Vision-Language Model Inference with Dynamical Text Guidance
- DreamVE: Unified Instruction-based Image and Video Editing
- Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- mKG-RAG: Multimodal Knowledge Graph-Enhanced RAG for Visual Question Answering
- VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
- IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection
- QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- VER-Bench: Evaluating MLLMs on Reasoning with Fine-Grained Visual Evidence
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- Training-Free Multimodal Large Language Model Orchestration
- Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
- Analyzing and Mitigating Object Hallucination: A Training Bias Perspective
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- TSPO: Temporal Sampling Policy Optimization for Long-form Video Language Understanding
- Enhancing Vision-Language Model Training with Reinforcement Learning in Synthetic Worlds for Real-World Success
- Gather and Trace: Rethinking Video TextVQA from an Instance-oriented Perspective
- Unlocking the Potential of MLLMs in Referring Expression Segmentation via a Light-weight Mask Decoder
- HPSv3: Towards Wide-Spectrum Human Preference Score
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- NaviMaster: Learning a Unified Policy for GUI and Embodied Navigation Tasks
- Accurate and Interpretable Postmenstrual Age Prediction via Multimodal Large Language Model
- StreamAgent: Towards Anticipatory Agents for Streaming Video Understanding
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- FluidFormer: Transformer with Continuous Convolution for Particle-based Fluid Simulation
- What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- Fine-grained Spatiotemporal Grounding on Egocentric Videos
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- UAV-ON: A Benchmark for Open-World Object Goal Navigation with Aerial Agents
- Oedipus and the Sphinx: Benchmarking and Improving Visual Language Models for Complex Graphic Reasoning
- PixNerd: Pixel Neural Field Diffusion
- Adversarial-Guided Diffusion for Multimodal LLM Attacks
- On the Risk of Misleading Reports: Diagnosing Textual Biases in Multimodal Clinical AI
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
- Extracting Visual Facts from Intermediate Layers for Mitigating Hallucinations in Multimodal Large Language Models
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
- BigTokDetect: A Clinically-Informed Vision-Language Modeling Framework for Detecting Pro-Bigorexia Videos on TikTok
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
- On the Reliability of Vision-Language Models Under Adversarial Frequency-Domain Perturbations
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding
- Secure Tug-of-War (SecTOW): Iterative Defense-Attack Training with Reinforcement Learning for Multimodal Model Security
- See Different, Think Better: Visual Variations Mitigating Hallucinations in LVLMs
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- Few-Shot Vision-Language Reasoning for Satellite Imagery via Verifiable Rewards
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- SafeDriveRAG: Towards Safe Autonomous Driving with Knowledge Graph-based Retrieval-Augmented Generation
- PRISM: Programmatic Reasoning with Image Sequence Manipulation for LVLM Jailbreaking
- MoHoBench: Assessing Honesty of Multimodal Large Language Models via Unanswerable Visual Questions
- Implicit Counterfactual Learning for Audio-Visual Segmentation
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems
- SessionIntentBench: A Multi-task Inter-session Intention-shift Modeling Benchmark for E-commerce Customer Behavior Understanding
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- FreeCus: Free Lunch Subject-driven Customization in Diffusion Transformers
- A Survey on Generative Model Unlearning: Fundamentals, Taxonomy, Evaluation, and Future Direction
- LAVA: Language Driven Scalable and Versatile Traffic Video Analytics
- Salsa as a Nonverbal Embodied Language -- The CoMPAS3D Dataset and Benchmarks
- Object-centric Video Question Answering with Visual Grounding and Referring
- A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation
- A Survey of Multimodal Hallucination Evaluation and Detection
- LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
- True Multimodal In-Context Learning Needs Attention to the Visual Context
- Advancing Complex Video Object Segmentation via Progressive Concept Construction
- EH-Benchmark Ophthalmic Hallucination Benchmark and Agent-Driven Top-Down Traceable Reasoning Workflow
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- ViGText: Deepfake Image Detection with Vision-Language Model Explanations and Graph Neural Networks
- VideoMind: An Omni-Modal Video Dataset with Intent Grounding for Deep-Cognitive Video Understanding
- Language Integration in Fine-Tuning Multimodal Large Language Models for Image-Based Regression
- Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark
- InterAct-Video: Reasoning-Rich Video QA for Urban Traffic
- VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization
- Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
- Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
- VLA-Mark: A cross modal watermark for large vision-language alignment model
- PIG-Nav: Key Insights for Pretrained Image Goal Navigation Models
- Eyes Will Shut: A Vision-Based Next GPS Location Prediction Model by Reinforcement Learning from Visual Map Feed Back
- Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
- Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
- Evaluating the Effectiveness of Cost-Efficient Large Language Models in Benchmark Biomedical Tasks
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
- VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
- Mitigating Object Hallucinations via Sentence-Level Early Intervention
- Evaluating Multimodal Large Language Models on Video Captioning via Monte Carlo Tree Search
- Human-like object concept representations emerge naturally in multimodal large language models
- POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
- VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism
- ReLoop-UME: Recurrent Depth with Learnable Retrieval Registers for Universal Multimodal Embedding
- CogDDN: A Cognitive Demand-Driven Navigation with Decision Optimization and Dual-Process Thinking
- NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking QA over Scientific Papers
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- IGD: Instructional Graphic Design with Multimodal Layer Generation
- A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
- Demonstrating the Octopi-1.5 Visual-Tactile-Language Model
- ElasticMM: Efficient Multimodal LLMs Serving with Elastic Multimodal Parallelism
- Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
- MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
- LaSM: Layer-wise Scaling Mechanism for Defending Pop-up Attack on GUI Agents
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
- VRU-Accident: A Vision-Language Benchmark for Video Question Answering and Dense Captioning for Accident Scene Understanding
- InsurTech innovation using natural language processing
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
- CAIRe: Cultural Attribution of Images by Retrieval-Augmented Evaluation
- SafeCoT: Improving VLM Safety with Minimal Reasoning
- Learning Diffusion Models with Flexible Representation Guidance
- Lumos-1: On Autoregressive Video Generation from a Unified Model Perspective
- Multilingual Multimodal Software Developer for Code Generation
- LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning
- A document is worth a structured record: Principled inductive bias design for document recognition
- Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
- VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models
- BlindSight: Harnessing Sparsity for Efficient Vision-Language Models
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
- MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- Scaling RL to Long Videos
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation
- VisualTrap: A Stealthy Backdoor Attack on GUI Agents via Visual Grounding Manipulation
- GR-LLMs: Recent Advances in Generative Recommendation Based on Large Language Models
- Omni-Video: Democratizing Unified Video Understanding and Generation
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- A Satellite-Ground Synergistic Large Vision-Language Model System for Earth Observation
- Skywork-R1V3 Technical Report
- TalkFashion: Intelligent Virtual Try-On Assistant Based on Multimodal Large Language Model
- SERUM: State Extraction and Refinement for User Modeling
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Transcribing Spanish Texts from the Past: Experiments with Transkribus, Tesseract and Granite
- From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
- Training-free Generation of Temporally Consistent Rewards from VLMs
- From Imitation to Innovation: The Emergence of AI Unique Artistic Styles and the Challenge of Copyright Protection
- Vision-Language Models Can't See the Obvious
- Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
- Forwardrobe: Garment-Aware Gaussian Avatars from a Single Image
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- Llama Nemoretriever Colembed: Top-Performing Text-Image Retrieval Model
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- ArtifactsBench: Bridging the Visual-Interactive Gap in LLM Code Generation Evaluation
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
- Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
- PresentAgent: Multimodal Agent for Presentation Video Generation
- LLM4Hint: Leveraging Large Language Models for Hint Recommendation in Offline Query Optimization
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- ZeroR@CHiPSAL 2026: Two-Stage Vision-Language Adaptation with Contrastive Learning for Nepali Meme Classification
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- No time to train! Training-Free Reference-Based Instance Segmentation
- From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
- LaCo: Efficient Layer-wise Compression of Visual Tokens for Multimodal Large Language Models
- TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- Look-Back: Implicit Visual Re-focusing in MLLM Reasoning
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
- Activation Reward Models for Few-Shot Model Alignment
- AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- SCING:Towards More Efficient and Robust Person Re-Identification through Selective Cross-modal Prompt Tuning
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
- Just Noticeable Difference for Large Multimodal Models
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- What Limits Virtual Agent Application? OmniBench: A Scalable Multi-Dimensional Benchmark for Essential Virtual Agent Capabilities
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- Token Activation Map to Visually Explain Multimodal LLMs
- VisualPrompter: Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
- Decoding Memes: Benchmarking Narrative Role Classification across Multilingual and Multimodal Models
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
- CyberV: Cybernetics for Test-time Scaling in Video Understanding
- UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
- Ovis-U1 Technical Report
- ActAlign: Zero-Shot Fine-Grained Video Classification via Language-Guided Sequence Alignment
- Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- Test-Time Consistency in Vision Language Models
- Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
- Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
- Universal Retrieval for Multimodal Trajectory Modeling
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
- MDC-R: The Minecraft Dialogue Corpus with Reference
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- Class-Agnostic Region-of-Interest Matching in Document Images
- LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
- Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- Agent-RewardBench: Towards a Unified Benchmark for Reward Modeling across Perception, Planning, and Safety in Real-World Multimodal Agents
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
- LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
- Mobile-R1: Towards Interactive Reinforcement Learning for VLM-Based Mobile Agent via Task-Level Rewards
- OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Recurrent Visual Feature Extraction and Stereo Attentions for CT Report Generation
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
- CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
- Unified Vision-Language-Action Model
- Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
- MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration
- Capturing Fine-Grained Alignments Improves 3D Affordance Detection
- Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
- Reading Smiles: Proxy Bias in Foundation Models for Facial Emotion Recognition
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
- Object-aware Sound Source Localization via Audio-Visual Scene Understanding
- Generalizing vision-language models to novel domains: A comprehensive survey
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation
- MUPA: Towards Multi-Path Agentic Reasoning for Grounded Video Question Answering
- CLGRPO: Reasoning Ability Enhancement for Small VLMs
- Adapting Vision-Language Models for Evaluating World Models
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
- PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- MDSAM:Memory-Driven Sparse Attention Matrix for LVLMs Hallucination Mitigation
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
- Hunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Condition
- WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
- The Role of Model Confidence on Bias Effects in Measured Uncertainties for Vision-Language Models
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Open World Scene Graph Generation using Vision Language Models
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- Cross-Modal Obfuscation for Jailbreak Attacks on Large Vision-Language Models
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- Structured Attention Matters to Multimodal LLMs in Document Understanding
- Visual symbolic mechanisms: Emergent symbol processing in vision language models
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Aligning Text, Images, and 3D Structure Token-by-Token
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- Understanding GUI Agent Localization Biases through Logit Sharpness
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- QUEST: Quality-aware Semi-supervised Table Extraction for Business Documents
- Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
- GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
- Dense360: Dense Understanding from Omnidirectional Panoramas
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
- EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
- MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
- Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- HKD4VLM: A Progressive Hybrid Knowledge Distillation Framework for Robust Multimodal Hallucination and Factuality Detection in VLMs
- State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
- VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
- GreedyPrune: Retenting Critical Visual Token Set for Large Vision Language Models
- RealHiTBench: A Comprehensive Realistic Hierarchical Table Benchmark for Evaluating LLM-Based Table Analysis
- Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models
- Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs
- MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks
- Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- LaTtE-Flow: Layerwise Timestep-Expert Flow-based Transformer
- The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
- Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
- Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
- Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025
- FlexRAG: A Flexible and Comprehensive Framework for Retrieval-Augmented Generation
- Mitigating Behavioral Hallucination in Multimodal Large Language Models for Sequential Images
- Are Multimodal Large Language Models Pragmatically Competent Listeners in Simple Reference Resolution Tasks?
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Representation Decomposition for Learning Similarity and Contrastness Across Modalities for Affective Computing
- VGR: Visual Grounded Reasoning
- Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
- Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
- EQA-RM: A Generative Embodied Reward Model with Test-time Scaling
- Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
- Burn After Reading: Do Multimodal Large Language Models Truly Capture Order of Events in Image Sequences?
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- DiMo-GUI: Advancing Test-time Scaling in GUI Grounding via Modality-Aware Visual Reasoning
- CogStream: Context-guided Streaming Video Question Answering
- VideoExplorer: Think With Videos For Agentic Long-Video Understanding
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
- LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities
- Challenging Vision-Language Models with Surgical Data: A New Dataset and Broad Benchmarking Study
- MadaKV: Adaptive Modality-Perception KV Cache Eviction for Efficient Multimodal Long-Context Inference
- CoMemo: LVLMs Need Image Context with Image Memory
- GenIR: Generative Visual Feedback for Mental Image Retrieval
- Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging
- MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- Unleashing Hour-Scale Video Training for Long Video-Language Understanding
- Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
- Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques
- MMRefine: Unveiling the Obstacles to Robust Refinement in Multimodal Large Language Models
- CIVET: Systematic Evaluation of Understanding in VLMs
- LLMs Can Compensate for Deficiencies in Visual Representations
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
- Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
- APVR: Hour-Level Long Video Understanding with Adaptive Pivot Visual Information Retrieval
- SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- REVEAL: Multi-turn Evaluation of Image-Input Harms for Vision LLM
- Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- MANBench: Is Your Multimodal Model Smarter than Human?
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
- DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
- LogisticsVLN: Vision-Language Navigation For Low-Altitude Terminal Delivery Based on Agentic UAVs
- Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- DFBench: Benchmarking Deepfake Image Detection Capability of Large Multimodal Models
- FlySearch: Exploring how vision-language models explore
- CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Seeing the Arrow of Time in Large Multimodal Models
- METok: Multi-Stage Event-based Token Compression for Efficient Long Video Understanding
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
- MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
- Contra4: Evaluating Contrastive Cross-Modal Reasoning in Audio, Video, Image, and 3D
- QARI-OCR: High-Fidelity Arabic Text Recognition through Multimodal Large Language Model Adaptation
- Flow2Code: Evaluating Large Language Models for Flowchart-based Code Generation Capability
- MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
- Fast-in-Slow: A Dual-System Foundation Model Unifying Fast Manipulation within Slow Reasoning
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- Respond Beyond Language: A Benchmark for Video Generation in Response to Realistic User Intents
- VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
- Follow the Flow: Fine-grained Flowchart Attribution with Neurosymbolic Agents
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- Is Extending Modality The Right Path Towards Omni-Modality?
- Generate, Not Recommend: Personalized Multimodal Content Generation
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- GOBench: Benchmarking Geometric Optics Generation and Understanding of MLLMs
- anyECG-chat: A Generalist ECG-MLLM for Flexible ECG Input and Multi-Task Understanding
- Aligning VLM Assistants with Personalized Situated Cognition
- Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times
- Improve MLLM Benchmark Efficiency through Interview
- Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
- Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
- GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
- WorldGym: World Model as An Environment for Policy Evaluation
- CReFT-CAD: Boosting Orthographic Projection Reasoning for CAD via Reinforcement Fine-Tuning
- The Security Threat of Compressed Projectors in Large Vision-Language Models
- Vid2Coach: Transforming How-To Videos into Task Assistants
- Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- Seeing the Abstract: Translating the Abstract Language for Vision Language Models
- Seg2Any: Open-set Segmentation-Mask-to-Image Generation with Precise Shape and Semantic Control
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
- BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
- CSVQA: A Chinese Multimodal Benchmark for Evaluating STEM Reasoning Capabilities of VLMs
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Who is in the Spotlight: The Hidden Bias Undermining Multimodal Retrieval-Augmented Generation
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
- FinMME: Benchmark Dataset for Financial Multi-Modal Reasoning Evaluation
- ClinBench-HPB: A Clinical Benchmark for Evaluating LLMs in Hepato-Pancreato-Biliary Diseases
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
- Time Blindness: Why Video-Language Models Can't See What Humans Can?
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
- ZeShot-VQA: Zero-Shot Visual Question Answering Framework with Answer Mapping for Natural Disaster Damage Assessment
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- Grounded Reinforcement Learning for Visual Reasoning
- OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
- VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
- Dataset Cartography for Large Language Model Alignment: Mapping and Diagnosing Preference Data
- SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- Are MLMs Trapped in the Visual Room?
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
- Fooling the Watchers: Breaking AIGC Detectors via Semantic Prompt Attacks
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- MMBoundary: Advancing MLLM Knowledge Boundary Awareness through Reasoning Step Confidence Calibration
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- Interpreting Chest X-rays Like a Radiologist: A Benchmark with Clinical Reasoning
- mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation
- DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- VidText: Towards Comprehensive Evaluation for Video Text Understanding
- Zero-Shot Vision Encoder Grafting via LLM Surrogates
- Scaling-up Perceptual Video Quality Assessment
- CADReview: Automatically Reviewing CAD Programs with Error Detection and Correction
- GETReason: Enhancing Image Context Extraction through Hierarchical Multi-Agent Reasoning
- HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- ChatVLA-2: Vision-Language-Action Model with Open-World Embodied Reasoning from Pretrained Knowledge
- Zero-Shot 3D Visual Grounding from Vision-Language Models
- cadrille: Multi-modal CAD Reconstruction with Reinforcement Learning
- Investigating Mechanisms for In-Context Vision Language Binding
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
- RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
- VIRAL: Vision-grounded Integration for Reward design And Learning
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
- Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
- DocReRank: Single-Page Hard Negative Query Generation for Training Multi-Modal RAG Rerankers
- Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
- Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
- SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
- NegVQA: Can Vision Language Models Understand Negation?
- MMTBENCH: A Unified Benchmark for Complex Multimodal Table Reasoning
- Leveraging Large Language Models in Visual Speech Recognition: Model Scaling, Context-Aware Decoding, and Iterative Polishing
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- XBOUND: Exploring Capability Boundaries of Device-Control Agents at the State Level
- On VLMs for Diverse Tasks in Multimodal Meme Classification
- Photography Perspective Composition: Towards Aesthetic Perspective Recommendation
- VAEmo: Efficient Representation Learning for Visual-Audio Emotion with Knowledge Injection
- TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
- BacktrackAgent: Enhancing GUI Agent with Error Detection and Backtracking Mechanism
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
- DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
- Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction
- MagicTryOn: Harnessing Diffusion Transformer for Garment-Preserving Video Virtual Try-on
- ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
- DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
- HoliTom: Holistic Token Merging for Fast Video Large Language Models
- LifeIR at the NTCIR-18 Lifelog-6 Task
- Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
- Predicting Implicit Arguments in Procedural Video Instructions
- CROP: Contextual Region-Oriented Visual Token Pruning
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
- Fast and Cost-effective Speculative Edge-Cloud Decoding with Early Exits
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
- MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
- OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
- AdaTP: Attention-Debiased Token Pruning for Video Large Language Models
- ReasonPlan: Unified Scene Prediction and Decision Reasoning for Closed-loop Autonomous Driving
- Attention! Your Vision Language Model Could Be Maliciously Manipulated
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
- Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
- Unveiling the Compositional Ability Gap in Vision-Language Reasoning Model
- DreamPRM: Domain-Reweighted Process Reward Model for Multimodal Reasoning
- CardioCoT: Hierarchical Reasoning for Multimodal Survival Analysis
- Jodi: Unification of Visual Generation and Understanding via Joint Modeling
- Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
- Sparse-to-Dense: A Free Lunch for Lossless Acceleration of Video Understanding in LLMs
- RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
- CCHall: A Novel Benchmark for Joint Cross-Lingual and Cross-Modal Hallucinations Detection in Large Language Models
- SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- MLLMs are Deeply Affected by Modality Bias
- ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- Inference Compute-Optimal Video Vision Language Models
- DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
- A General-Purpose VLM Can Teach an Astronomy Foundation Model to Better Recognize Galaxy Morphology
- Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations
- FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
- Towards General Continuous Memory for Vision-Language Models
- TransBench: Breaking Barriers for Transferable Graphical User Interface Agents in Dynamic Digital Environments
- Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports
- Scaling Image and Video Generation via Test-Time Evolutionary Search
- Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
- Illuminating Visual Identity in Universal Multimodal Embeddings
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in the Financial Domain
- Diagnosing Vision Language Models' Perception by Leveraging Human Methods for Color Vision Deficiencies
- Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
- GROVE: Growing and Reasoning over Temporally Stratified Memory from Streaming Video Experience
- MDIT-Bench: Evaluating the Dual-Implicit Toxicity in Large Multimodal Models
- SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
- Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
- Benchmarking Retrieval-Augmented Multimodal Generation for Document Question Answering
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- Training-Free Reasoning and Reflection in MLLMs
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- GUI-explorer: Autonomous Exploration and Mining of Transition-aware Knowledge for GUI Agent
- Panoptic Captioning: An Equivalence Bridge for Image and Text
- Recursive Offloading for LLM Serving in Multi-tier Networks
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
- Hierarchical Safety Realignment: Lightweight Restoration of Safety in Pruned Large Vision-Language Models
- SoccerChat: Integrating Multimodal Data for Enhanced Soccer Game Understanding
- MMaDA: Multimodal Large Diffusion Language Models
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
- ChartCards: A Chart-Metadata Generation Framework for Multi-Task Chart Understanding
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
- Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
- Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
- Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
- Dynamic Resolution Routing for Efficient Egocentric Grounding
- SNAP: A Benchmark for Testing the Effects of Capture Conditions on Fundamental Vision Tasks
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Decouple and Orthogonalize: A Data-Free Framework for LoRA Merging
- TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models
- OViP: Online Vision-Language Preference Learning for VLM Hallucination
- Clapper: Compact Learning and Video Representation in VLMs
- How Do Large Vision-Language Models See Text in Image? Unveiling the Distinctive Role of OCR Heads
- CRAFT: Compression via Recursive Adaptive Fusion of Video Tokens for Vision-Language Models
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question Answering
- Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
- Emerging Properties in Unified Multimodal Pretraining
- Inter-Residue Geometry Attention for Antibody-Specific Epitope Prediction
- Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?
- Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
- VoQA: Visual-only Question Answering
- Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
- Dual Precision Quantization for Efficient and Accurate Deep Neural Networks Inference
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
- Multimodal RAG-driven Anomaly Detection and Classification in Laser Powder Bed Fusion using Large Language Models
- Visual Agentic Reinforcement Fine-Tuning
- LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- Debating for Better Reasoning: An Unsupervised Multimodal Approach
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models
- ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method
- Memory-Centric Embodied Question Answering
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Texts or Images? A Fine-grained Analysis on the Effectiveness of Input Representations and Models for Table Question Answering
- Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
- PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models
- Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis
- 3D Visual Illusion Depth Estimation
- Specialized Foundation Models for Intelligent Operating Rooms
- Reasoning-OCR: Can Large Multimodal Models Solve Complex Logical Reasoning Problems from OCR Cues?
- HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
- GeoRanker: Distance-Aware Ranking for Worldwide Image Geolocalization
- FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
- STAR-VLM: Spatiotemporal Grounding Vision-Language Models for Motion and Velocity Estimation via Automotive Radar Supervision
- Advancing Sequential Numerical Prediction in Autoregressive Models
- MobileIPL: Enhancing Mobile Agents Thinking Process via Iterative Preference Learning
- NeuroGen: Neural Network Parameter Generation via Large Language Models
- CompBench: Benchmarking Complex Instruction-guided Image Editing
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- TinyRS-R1: Compact Multimodal Language Model for Remote Sensing
- PRS-Med: Position Reasoning Segmentation in Medical Imaging
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- Mobile-Bench-v2: A More Realistic and Comprehensive Benchmark for VLM-based Mobile Agents
- OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Are vision language models robust to uncertain inputs?
- UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings
- ChartEdit: How Far Are MLLMs From Automating Chart Analysis? Evaluating MLLMs' Capability via Chart Editing
- Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction
- TCC-Bench: Benchmarking the Traditional Chinese Culture Understanding Capabilities of MLLMs
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
- HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
- Cross-Image Contrastive Decoding: Precise, Lossless Suppression of Language Priors in Large Vision-Language Models
- Decoding the Multimodal Mind: Generalizable Brain-to-Text Translation via Multimodal Alignment and Adaptive Routing
- Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
- Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
- PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
- MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning
- Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
- MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
- VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
- Flash-VL 2B: Optimizing Vision-Language Model Performance for Ultra-Low Latency and High Throughput
- Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
- ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
- Behind Maya: Building a Multilingual Vision Language Model
- Advancing Food Nutrition Estimation via Visual-Ingredient Feature Fusion
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- DeceptionX: From Multimodal Evidence to Explainable Deception Detection
- HorusEye: Language as Dynamic Attention for Emergency Visual Analysis
- Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
- Ultrasound Report Generation with Multimodal Large Language Models for Standardized Texts
- Large Language Models for Computer-Aided Design: A Survey
- Visually Interpretable Subtask Reasoning for Visual Question Answering
- SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing
- EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
- DanceGRPO: Unleashing GRPO on Visual Generation
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- DocVXQA: Context-Aware Visual Explanations for Document Question Answering
- EgoIntent: A Pre-Outcome Micro-Step Benchmark for Understanding What, Why, and Next
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
- Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
- Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
- GPrune-LLM: Generalization-Aware Structured Pruning for Large Language Models
- Emotion-Qwen: A Unified Framework for Emotion and Vision Understanding
- Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
- An Analysis Focused on Womens Safety: Can VAD Models Be Enhanced by a Multi-modal Dataset?
- Integrating Video and Text: A Balanced Approach to Multimodal Summary Generation and Evaluation
- ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models
- VISTA: Generative Visual Imagination for Vision-and-Language Navigation
- Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
- Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
- Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
- FG-CLIP: Fine-Grained Visual and Textual Alignment
- Hierarchical Pre-Training of Vision Encoders with Large Language Model
- MMPhysVideo: Physically Plausible Video Generation Through Joint RGB-Perception Modeling
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
- UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
- MateICL: Mitigating Attention Dispersion in Large-Scale In-Context Learning
- Efficient Vocabulary-Free Fine-Grained Visual Recognition in the Age of Multimodal LLMs
- Controllable Weather Synthesis and Removal with Video Diffusion Models
- Would you still call this Dax? Novel Visual References in VLMs and Humans
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- In-Context Collapse in Vision-Language Models and How to Mitigate it?
- Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
- Reward Valuation in Vision Language Models: Causal Mechanisms Underlying Anhedonia
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
- ScaleTrack: Scaling and back-tracking Automated GUI Agents
- MAIN-VLA: Modeling Abstraction of Intention and eNvironment for Vision-Language-Action Models
- MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals
- FireRed-OCR Technical Report
- AVA: Towards Agentic Video Analytics with Vision Language Models
- Can Generalist Agents Automate Data Curation?
- Do VLMs Align Better with Humans than LLMs during Natural Reading?
- Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments
- Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
- Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering
- SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding
- Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
- Multimodal Language Models See Better When They Look Shallower
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- DEL: Digit Entropy Loss for Numerical Learning of Large Language Models
- GROW: Aligning GRPO with State-Action Modeling for Open-World VLM Agents
- Runtime Monitoring of Perception-Based Autonomous Systems via Embedding Temporal Logic
- Follow the Mean: Reference-Guided Flow Matching
- Computational Reasoning of Large Language Models
- LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
- FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
- Learning Streaming Video Representation via Multitask Training
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
- Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
- Let ViT Speak: Generative Language-Image Pre-training
- EEG-Based Brain-LLM Interface for Human Preference Aligned Generation
- Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer
- Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
- LR-IAD:Mask-Free Industrial Anomaly Detection with Logical Reasoning
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
- Reason Like a Radiologist: Chain-of-Thought and Reinforcement Learning for Verifiable Report Generation
- StarVLA-α: Reducing Complexity in Vision-Language-Action Systems
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- SMART: When is it Actually Worth Expanding a Speculative Tree?
- Visual and Textual Prompts in VLLMs for Enhancing Emotion Recognition
- Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
- Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion
- A Review of 3D Object Detection with Vision-Language Models
- Sky-Drive: A Distributed Multi-Agent Simulation Platform for Human-AI Collaborative and Socially-Aware Future Transportation
- Proof-of-TBI -- Fine-Tuned Vision Language Model Consortium and OpenAI-o3 Reasoning LLM-Based Medical Diagnosis Support System for Mild Traumatic Brain Injury (TBI) Prediction
- VEU-Bench: Towards Comprehensive Understanding of Video Editing
- DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
- Robotic Task Ambiguity Resolution via Natural Language Interaction
- Bayesian Data Reweighting Improves Multimodal Retrieval for Knowledge-Based Visual Question Answering
- TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
- A Model Merging Approach for Continual MLLM Unlearning
- REZE: Recognition-Based Zero-Shot Extraction for Video Temporal Grounding
- CARVE: Cross-Slice Anisotropic Reallocation of Visual Evidence for Efficient 3D Medical Volume Understanding
- DIVE: Dynamic Iterative Visual Evidence Construction for Efficient Vision-Language Models
- When Prompts Become Pixels: Prompt-Region Grounding for Multimodal Reasoning
- HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding
- Persistent Object Narratives for Token-Efficient Video Language Models
- RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
- GEB-Bench: Abstract Structures Told in Many Voices
- URECA: Unique Region Caption Anything
- Latent Denoising Improves Visual Alignment in Large Multimodal Models
- Towards Visual Text Grounding of Multimodal Large Language Model
- VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
- VideoVista-CulturalLingo: 360^∘ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
- On the Robustness of GUI Grounding Models Against Image Attacks
- Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- Unveiling the Lack of LVLM Robustness to Fundamental Visual Variations: Why and Path Forward
- Describe Anything: Detailed Localized Image and Video Captioning
- ViSMaP: Unsupervised Hour-long Video Summarisation by Meta-Prompting
- Advancing Egocentric Video Question Answering with Multimodal Large Language Models
- Vidi: Large Multimodal Models for Video Understanding and Editing
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
- Object-Level Verbalized Confidence Calibration in Vision-Language Models via Semantic Perturbation
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- Are Vision LLMs Road-Ready? A Comprehensive Benchmark for Safety-Critical Driving Video Understanding
- Learning from Reasoning Failures via Synthetic Data Generation
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- Enhancing Multimodal In-Context Learning for Image Classification through Coreset Optimization
- PipeWeaver: Addressing Data Dynamicity in Large Multimodal Model Training with Dynamic Interleaved Pipeline
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- LearnAct: Few-Shot Mobile GUI Agent with a Unified Demonstration Benchmark
- Visual Intention Grounding for Egocentric Assistants
- Compile Scene Graphs with Reinforcement Learning
- MusFlow: Multimodal Music Generation via Conditional Flow Matching
- ChartQA-X: Generating Explanations for Visual Chart Reasoning
- Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
- VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
- TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- Beyond Text: Characterizing Domain Expert Needs in Document Research
- Beauty and the Bias: Exploring the Impact of Attractiveness on Multimodal Large Language Models
- FLIP Reasoning Challenge
- Visual Grounding in Zero-Shot Vision-Language Control
- Coherence-Oriented Dream Scene Visualisation
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- HoloCount: A Holistic Visual Counting Benchmark for MLLMs
- TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
- VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
- RIG-RoPE: Relation- and Instance-Gated Rotary Positional Encoding with Duration-Aware Temporal Coordinates
- Look Twice: Training-Free Evidence Highlighting for Knowledge-based Visual Question Answering
- Never Start from Scratch: Expediting On-Device LLM Personalization via Explainable Model Selection
- PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
- PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving
- Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
- Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
- MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
- SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
- Relation-Rich Visual Document Generator for Visual Information Extraction
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- Summarization of Multimodal Presentations with Vision-Language Models: Study of the Effect of Modalities and Structure
- OVERLORD: Ultimate Scaling of DataLoader for Multi-Source Large Foundation Model Training
- Unchecked and Overlooked: Addressing the Checkbox Blind Spot in Large Language Models with CheckboxQA
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Aligning Anime Video Generation with Human Feedback
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
- SemiETS: Integrating Spatial and Content Consistencies for Semi-Supervised End-to-end Text Spotting
- Resampling Benchmark for Efficient Comprehensive Evaluation of Large Vision-Language Models
- Breaking the Data Barrier -- Building GUI Agents Through Task Generalization
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- Multimodal Long Video Modeling Based on Temporal Dynamic Context
- MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework
- AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
- GenTe: Generative Real-world Terrains for General Legged Robot Locomotion Control
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
- Distilling Transitional Pattern to Large Language Models for Multimodal Session-based Recommendation
- Don't Deceive Me: Mitigating Gaslighting through Attention Reallocation in LMMs
- TextSplat: Text-Guided Semantic Fusion for Generalizable Gaussian Splatting
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment
- VideoAds for Fast-Paced Video Understanding
- Towards Explainable Partial-AIGC Image Quality Assessment
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- Mixed Signals: Decoding VLMs' Reasoning and Underlying Bias in Vision-Language Conflict
- Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model
- Mimic In-Context Learning for Multimodal Tasks
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- VisualCloze: A Universal Image Generation Framework via Visual In-Context Learning
- MM-IFEngine: Towards Multimodal Instruction Following
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Scaling Laws for Native Multimodal Models
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- Classifying the Unknown: In-Context Learning for Open-Vocabulary Text and Symbol Recognition
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- OmniCaptioner: One Captioner to Rule Them All
- SCI-Reason: A Dataset with Chain-of-Thought Rationales for Complex Multimodal Reasoning in Academic Areas
- SafeMLRM: Demystifying Safety in Multi-modal Large Reasoning Models
- Distilling Textual Priors from LLM to Efficient Image Fusion
- V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
- Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
- The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
- ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering
- Rethinking RoPE: A Mathematical Blueprint for N-dimensional Positional Embedding
- SmolVLM: Redefining small and efficient multimodal models
Related