Visual Instruction Tuning
2023/04/17 by Haotian Liu, Liu, Haotian, Chunyuan Li +5 · 2,497 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques
paper · pdf · doi:10.48550/arxiv.2304.08485
openalex publication_date 2023/04/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/30
Abstract
Instruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field. In this paper, we present the first attempt to use language-only GPT-4 to generate multimodal language-image instruction-following data. By instruction tuning on such generated data, we introduce LLaVA: Large Language and Vision Assistant, an end-to-end trained large multimodal model that connects a vision encoder and LLM for general-purpose visual and language understanding.Our early experiments show that LLaVA demonstrates impressive multimodel chat abilities, sometimes exhibiting the behaviors of multimodal GPT-4 on unseen images/instructions, and yields a 85.1% relative score compared with GPT-4 on a synthetic multimodal instruction-following dataset. When fine-tuned on Science QA, the synergy of LLaVA and GPT-4 achieves a new state-of-the-art accuracy of 92.53%. We make GPT-4 generated visual instruction tuning data, our model and code base publicly available.
Cited by
- DeCoRAG: Cognitive Decoupling and Semantic-Aware Cropping for Complex Document Understanding
- Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning
- DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes
- UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- ProGuard: Towards Proactive Multimodal Safeguard
- Instruction-Following Evaluation of Large Vision-Language Models
- D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models
- CoFi-Dec: Hallucination-Resistant Decoding via Coarse-to-Fine Generative Feedback in Large Vision-Language Models
- Video Understanding: From Geometry and Semantics to Unified Models
- SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
- An Architecture-Led Hybrid Report on Body Language Detection Project
- Embodied Robot Manipulation in the Era of Foundation Models: Planning and Learning Perspectives
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- Meta-information Guided Cross-domain Synergistic Diffusion Model for Low-dose PET Reconstruction
- VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- Gaze-to-text Generation: Beyond Categorical Decoding of Human Attention
- Towards High-Level Semantic Intelligence
- MemVLN: Episodic and Procedural Memory for Vision-and-Language Navigation
- RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
- Hybrid Advantage Estimation with Unified Critic for VLM Agentic Reinforcement Learning
- Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning
- PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- MEDit-Bench: A Dataset for Evaluating Message-Driven Narrative Video Editing
- Salient Knowledge Pathways: Sparse Cross-Modal Routing for Efficient Knowledge-Intensive Multimodal Question Answering
- Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering
- Noise-Free One-Step LoRA for Task-Driven Image Restoration with Diffusion Priors
- LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed
- LongFly: Long-Horizon UAV Vision-and-Language Navigation with Spatiotemporal Context Integration
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- Controlling Embedding Spaces with Text-Conditioned Transformations
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
- An LMM for Precisely Grounding Elements in Documents
- ReSAGE-PAR: Representational Similarity Assessment for Generative Expansion in Pedestrian Attribute Recognition
- Aligning Quantum Operators with Large Language Models
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- NEXT: Reasoning-Driven Video Recommendation via a Vision-Language Model
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
- Personalize Your Large Vision-language Models With In-context Prompt Tuning
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- LanteRn: Latent Visual Structured Reasoning
- StAR: Segment Anything Reasoner
- Compressing LLMs with MoP: Mixture of Pruners
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Training-free Conditional Image Embedding Framework Leveraging Large Vision Language Models
- ImagineNav++: Prompting Vision-Language Models as Embodied Navigator through Scene Imagination
- A Medical Multimodal Diagnostic Framework Integrating Vision-Language Models and Logic Tree Reasoning
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
- Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Streaming Video Instruction Tuning
- Fast SAM2 with Text-Driven Token Pruning
- Latent Implicit Visual Reasoning
- ORCA: Object Recognition and Comprehension for Archiving Marine Species
- MarineEval: Assessing the Marine Intelligence of Vision-Language Models
- Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
- PanoGrounder: Bridging 2D and 3D with Panoramic Scene Representations for VLM-based 3D Visual Grounding
- Input-Adaptive Visual Preprocessing for Efficient Fast Vision-Language Model Inference
- FlashVLM: Text-Guided Visual Token Selection for Large Multimodal Models
- Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
- M3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Generative Digital Twins: Vision-Language Simulation Models for Executable Industrial Systems
- SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- Event Extraction in Large Language Model
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- ReasonCD: A Multimodal Reasoning Large Model for Implicit Change-of-Interest Semantic Mining
- Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- Delta-LLaVA: Base-then-Specialize Alignment for Token-Efficient Vision-Language Models
- Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- EchoMotion: Unified Human Video and Motion Generation via Dual-Modality Diffusion Transformer
- Revealing Perception and Generation Dynamics in LVLMs: Mitigating Hallucinations via Validated Dominance Correction
- LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
- M3-Verse: A "Spot the Difference" Challenge for Large Multimodal Models
- OpenView: Empowering MLLMs with Out-of-view VQA
- Enhancing 3D Semantic Scene Completion with a Refinement Module
- MoE Pathfinder: Trajectory-driven Expert Pruning
- HyDRA: Hierarchical and Dynamic Rank Adaptation for Mobile Vision Language Model
- SLIM: Semantic-based Low-bitrate Image compression for Machines by leveraging diffusion
- External Hippocampus: Topological Cognitive Maps for Guiding Large Language Model Reasoning
- FPBench: A Comprehensive Benchmark of Multimodal Large Language Models for Fingerprint Analysis
- Robotic VLA Benefits from Joint Learning with Motion Image Diffusion
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- AnyTask: an Automated Task and Data Generation Framework for Advancing Sim-to-Real Policy Learning
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Xiaomi MiMo-VL-Miloco Technical Report
- RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
- Are Vision Language Models Cross-Cultural Theory of Mind Reasoners?
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Vision-Language Model Guided Image Restoration
- CheXPO-v2: Preference Optimization for Chest X-ray VLMs with Knowledge Graph Consistency
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- Monte Carlo Query Search: Active Capability Assessment of AI Agents
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
- Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment
- AMUSE: Audio-Visual Benchmark and Alignment Framework for Agentic Multi-Speaker Understanding
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- Sceniris: A Fast Procedural Scene Generation Framework
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Large Video Planner Enables Generalizable Robot Control
- An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
- VLA-AN: An Efficient and Onboard Vision-Language-Action Framework for Aerial Navigation in Complex Environments
- On-device Large Multi-modal Agent for Human Activity Recognition
- EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- Unified Semantic Transformer for 3D Scene Understanding
- Vector Prism: Animating Vector Graphics by Stratifying Semantic Structure
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- Autonomous Construction-Site Safety Inspection Using Mobile Robots: A Multilayer VLM-LLM Pipeline
- SuperCLIP: CLIP with Simple Classification Supervision
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- Beyond surface form: A pipeline for semantic analysis in Alzheimer's Disease detection from spontaneous speech
- AgentIAD: Agentic Industrial Anomaly Detection via Adaptive Memory Augmentation
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- Reflective Preference Optimization (RPO): Enhancing On-Policy Alignment via Hint-Guided Reflection
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- Why Text Prevails: Vision May Undermine Multimodal Medical Decision Making
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Efficient Vision-Language Reasoning via Adaptive Token Pruning
- Content-Aware Ad Banner Layout Generation with Two-Stage Chain-of-Thought in Vision Language Models
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- EditMGT: Unleashing Potentials of Masked Generative Transformers in Image Editing
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- VEGAS: Mitigating Hallucinations in Large Vision-Language Models via Vision-Encoder Attention Guided Adaptive Steering
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- HFS: Holistic Query-Aware Frame Selection for Efficient Video Reasoning
- Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
- DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
- VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- MeViS: A Multi-Modal Dataset for Referring Motion Expression Video Segmentation
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- BabyVLM-V2: Toward Developmentally Grounded Pretraining and Benchmarking of Vision Foundation Models
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
- Towards Fine-Grained Recognition with Large Visual Language Models: Benchmark and Optimization Strategies
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- Multilingual VLM Training: Adapting an English-Trained VLM to French
- Investigating The Functional Roles of Attention Heads in Vision Language Models: Evidence for Reasoning Modules
- Grounding Everything in Tokens for Multimodal Large Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- FlipLLM: Efficient Bit-Flip Attacks on Multimodal LLMs using Reinforcement Learning
- IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
- Investigate the Low-level Visual Perception in Vision-Language based Image Quality Assessment
- Building Reasonable Inference for Vision-Language Models in Blind Image Quality Assessment
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- Video-QTR: Query-Driven Temporal Reasoning Framework for Lightweight Video Understanding
- FoundIR-v2: Optimizing Pre-Training Data Mixtures for Image Restoration Foundation Model
- GLACIA: Instance-Aware Positional Reasoning for Glacial Lake Segmentation via Multimodal Large Language Model
- Explaining the Unseen: Multimodal Vision-Language Reasoning for Situational Awareness in Underground Mining Disasters
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- Towards Lossless Ultimate Vision Token Compression for VLMs
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- PAVAS: Physics-Aware Video-to-Audio Synthesis
- HybridToken-VLM: Hybrid Token Compression for Vision-Language Models
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- Embodied Tree of Thoughts: Deliberate Manipulation Planning with Embodied World Model
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- Instance-Aware Test-Time Segmentation for Continual Domain Shifts
- LUNA: Linear Universal Neural Attention with Generalization Guarantees
- Relational Visual Similarity
- One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
- HalluShift++: Bridging Language and Vision through Internal Representation Shifts for Hierarchical Hallucinations in MLLMs
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- All You Need Are Random Visual Tokens? Demystifying Token Pruning in VLLMs
- Toward More Reliable Artificial Intelligence: Reducing Hallucinations in Vision-Language Models
- Zero-Shot Textual Explanations via Translating Decision-Critical Features
- Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
- START: Spatial and Textual Learning for Chart Understanding
- SPACE: Noise Contrastive Estimation Stabilizes Self-Play Fine-Tuning for Large Language Models
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information Interaction
- NeuroABench: A Multimodal Evaluation Benchmark for Neurosurgical Anatomy Identification
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- The Role of Entropy in Visual Grounding: Analysis and Optimization
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Scaling Zero-Shot Reference-to-Video Generation
- Personalized Image Descriptions from Attention Sequences
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- BeLLA: End-to-End Birds Eye View Large Language Assistant for Autonomous Driving
- ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior
- Evaluating Concept Filtering Defenses against Child Sexual Abuse Material Generation by Text-to-Image Models
- 2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency
- Concept-based Explainable Data Mining with VLM for 3D Detection
- LoC-Path: Learning to Compress for Pathology Multimodal Large Language Models
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- FASTer: Toward Efficient Autoregressive Vision Language Action Modeling via Neural Action Tokenization
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- Qwen3.5-Omni Technical Report
- OmniScaleSR: Unleashing Scale-Controlled Diffusion Prior for Faithful and Realistic Arbitrary-Scale Image Super-Resolution
- Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- BiTAgent: A Task-Aware Modular Framework for Bidirectional Coupling between Multimodal Large Language Models and World Models
- Jina-VLM: Small Multilingual Vision Language Model
- C3G: Learning Compact 3D Representations with 2K Gaussians
- DIQ-H: Evaluating Hallucination Persistence in VLMs Under Temporal Visual Degradation
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- Tutorial on Large Language Model-Enhanced Reinforcement Learning for Wireless Networks
- Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- YOLOA: Real-Time Affordance Detection via LLM Adapter
- Large Language Models as Generalist Policies for Network Optimization
- Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- PPTBench: Towards Holistic Evaluation of Large Language Models for PowerPoint Layout and Design Understanding
- PGP-DiffSR: Phase-Guided Progressive Pruning for Efficient Diffusion-based Image Super-Resolution
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
- See, Think, Learn: A Self-Taught Multimodal Reasoner
- VACoT: Rethinking Visual Data Augmentation with VLMs
- Understanding and Harnessing Sparsity in Unified Multimodal Models
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- Progressive Image Restoration via Text-Conditioned Video Generation
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- LLM CHESS: Benchmarking Reasoning and Instruction-Following in LLMs through Chess
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- Med-VCD: Mitigating Hallucination for Medical Large Vision Language Models through Visual Contrastive Decoding
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- Generative Action Tell-Tales: Assessing Human Motion in Synthesized Videos
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- LEC: Linear Expectation Constraints for Selection-Conditioned Risk Control in Selective Prediction and Routing Systems
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- MCAT: Scaling Many-to-Many Speech-to-Text Translation with MLLMs to 70 Languages
- S2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
- SocialDriveGen: Generating Diverse Traffic Scenarios with Controllable Social Interactions
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
- DefenSee: Dissecting Threat from Sight and Text -- A Multi-View Defensive Pipeline for Multi-modal Jailbreaks
- KidSpeak: A General Multi-purpose LLM for Kids' Speech Recognition and Screening
- Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
- LLM2Fx-Tools: Tool Calling For Music Post-Production
- See, Hear, and Understand: Benchmarking Audiovisual Human Speech Understanding in Multimodal Large Language Models
- LISA-3D: Lifting Language-Image Segmentation to 3D via Multi-View Consistency
- PhotoFramer: Multi-modal Image Composition Instruction
- Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
- Table as a Modality for Large Language Models
- AFRAgent : An Adaptive Feature Renormalization Based High Resolution Aware GUI agent
- Minimal neuron ablation triggers catastrophic collapse in the language core of Large Vision-Language Models
- Multilingual Training-Free Remote Sensing Image Captioning
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- BioPro: Towards Difference-Aware Gender Fairness for Vision-Language Models
- DEJIMA: A Novel Large-scale Japanese Dataset for Image Captioning and Visual Question Answering
- Concept-Guided Backdoor Attack on Vision Language Models
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Low-Bitrate Video Compression through Semantic-Conditioned Diffusion
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- BioArc: Discovering Optimal Neural Architectures for Biological Foundation Models
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- DenseScan: Advancing 3D Scene Understanding with 2D Dense Annotation
- Mammo-FM: Breast-specific foundational model for Integrated Mammographic Diagnosis, Prognosis, and Reporting
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- DEAL-300K: Diffusion-based Editing Area Localization with a 300K-Scale Dataset and Frequency-Prompted Baseline
- Optimizing Multimodal Language Models through Attention-based Interpretability
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- Buffer replay enhances the robustness of multimodal learning under missing-modality
- HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
- Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
- Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
- Revisiting the Necessity of Lengthy Chain-of-Thought in Vision-centric Reasoning Generalization
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- Advancing Aesthetic Image Generation via Composition Transfer
- Beyond Real versus Fake Towards Intent-Aware Video Analysis
- GEO-Detective: Unveiling Location Privacy Risks in Images with LLM Agents
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- Unexplored flaws in multiple-choice VQA evaluations
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- Exploring Diagnostic Prompting Approach for Multimodal LLM-based Visual Complexity Assessment: A Case Study of Amazon Search Result Pages
- TrafficLens: Multi-Camera Traffic Video Analysis Using LLMs
- Towards Reasoning-Preserving Unlearning in Multimodal Large Language Models
- DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies
- Text-Guided Semantic Image Encoder
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
- Object-Centric Vision Token Pruning for Vision Language Models
- Thinking in 360°: Humanoid Visual Search in the Wild
- V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs
- Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- CoC-VLA: Delving into Adversarial Domain Transfer for Explainable Autonomous Driving via Chain-of-Causality Visual-Language-Action Model
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
- MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Fara-7B: An Efficient Agentic Model for Computer Use
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- HunyuanOCR Technical Report
- Evaluating Dataset Watermarking for Fine-tuning Traceability of Customized Diffusion Models: A Comprehensive Benchmark and Removal Approach
- Cross Domain Evaluation of Multimodal Chain-of-Thought Reasoning of different datasets into the Amazon CoT Framework
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Percept-WAM: Perception-Enhanced World-Awareness-Action Model for Robust End-to-End Autonomous Driving
- RAVEN++: Pinpointing Fine-Grained Violations in Advertisement Videos with Active Reinforcement Reasoning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- EEG-VLM: A Hierarchical Vision-Language Model with Multi-Level Feature Alignment and Visually Enhanced Language-Guided Reasoning for EEG Image-Based Sleep Stage Prediction
- Collaborative Learning with Multiple Foundation Models for Source-Free Domain Adaptation
- AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
- BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- Understanding Task Transfer in Vision-Language Models
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation
- Extreme Model Compression for Edge Vision-Language Models: Sparse Temporal Token Fusion and Adaptive Neural Compression
- DAVSP: Safety Alignment for Large Vision-Language Models via Deep Aligned Visual Safety Prompt
- SineProject: Machine Unlearning for Stable Vision Language Alignment
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
- MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
- Weakly-supervised Latent Models for Task-specific Visual-Language Control
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post-hoc Debiasing in Vision-Language Models
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- Consolidating Diffusion-Generated Video Detection with Unified Multimodal Forgery Learning
- Multi-speaker Attention Alignment for Multimodal Social Interaction
- SPINE: Token-Selective Test-Time Reinforcement Learning with Entropy-Band Regularization
- MobileVLA-R1: Reinforcing Vision-Language-Action for Mobile Robots
- VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
- Attention Guided Alignment in Efficient Vision-Language Models
- A cross-species neural foundation model for end-to-end speech decoding
- CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Q-MLLM: Vector Quantization for Robust Multimodal Large Language Model Security
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- Bridging the Semantic-Action Gap in Visual Token Pruning for Efficient VLA Inference
- Do Vision-Language Models Understand Visual Persuasiveness?
- UAM: A Unified Attention-Mamba Backbone of Multimodal Framework for Tumor Cell Classification
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation
- TeamPath: Building MultiModal Pathology Experts with Reasoning AI Copilots
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- TS-PEFT: Unveiling Token-Level Redundancy in Parameter-Efficient Fine-Tuning
- LLMs-based Augmentation for Domain Adaptation in Long-tailed Food Datasets
- Boosting Medical Visual Understanding From Multi-Granular Language Learning
- An Image Is Worth Ten Thousand Words: Verbose-Text Induction Attacks on VLMs
- Fairness in Multi-modal Medical Diagnosis with Demonstration Selection
- Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
- Joint Semantic-Channel Coding and Modulation for Token Communications
- HV-Attack: Hierarchical Visual Attack for Multimodal Retrieval Augmented Generation
- Parameter Importance-Driven Continual Learning for Foundation Models
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- Instruction-Guided Lesion Segmentation for Chest X-rays with Automatically Generated Large-Scale Dataset
- Multimodal Continual Instruction Tuning with Dynamic Gradient Guidance
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Look, Zoom, Understand: The Robotic Eyeball for Embodied Perception
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
- SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
- Stealth Fine-Tuning: Efficiently Breaking Alignment in RVLMs Using Self-Generated CoT
- Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
- Weakly Supervised Ephemeral Gully Detection In Remote Sensing Images Using Vision Language Models
- Learning Skill-Attributes for Transferable Assessment in Video
- VLMs Guided Interpretable Decision Making for Autonomous Driving
- Unlocking the Forgery Detection Potential of Vanilla MLLMs: A Novel Training-Free Pipeline
- Moving Pictures of Thought: Extracting Visual Knowledge in Charles S. Peirce's Manuscripts with Vision-Language Models
- Dual-LoRA and Quality-Enhanced Pseudo Replay for Multimodal Continual Food Learning
- Video Finetuning Improves Reasoning Between Frames
- MMD-Thinker: Adaptive Multi-Dimensional Thinking for Multimodal Misinformation Detection
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Black-Box Membership Inference Attack for LVLMs via Prior Knowledge-Calibrated Memory Probing
- Direct Visual Grounding by Directing Attention of Visual Tokens
- FSDAM: Few-Shot Driving Attention Modeling via Vision-Language Coupling
- AERMANI-VLM: Structured Prompting and Reasoning for Aerial Manipulation with Vision Language Models
- HMVLM: Human Motion-Vision-Lanuage Model via MoE LoRA
- Can large language models be a cardinality estimator? An empirical study
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- CoTBox-TTT: Grounding Medical VQA with Visual Chain-of-Thought Boxes During Test-time Training
- SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- Fast Reasoning Segmentation for Images and Videos
- Suppressing VLM Hallucinations with Spectral Representation Filtering
- MediRound: Multi-Round Entity-Level Reasoning Segmentation in Medical Images
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- PAS : Prelim Attention Score for Detecting Object Hallucinations in Large Vision--Language Models
- EcoAlign: An Economically Rational Framework for Efficient LVLM Alignment
- MAFM3: Modular Adaptation of Foundation Models for Multi-Modal Medical AI
- Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Draft and Refine with Visual Experts
- Binary Verification for Zero-Shot Vision
- Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- TransactionGPT
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
- Anatomy-VLM: A Fine-grained Vision-Language Model for Medical Interpretation
- VectorSynth: Fine-Grained Satellite Image Synthesis with Structured Semantics
- Remodeling Semantic Relationships in Vision-Language Fine-Tuning
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- VADER: Towards Causal Video Anomaly Understanding with Relation-Aware Large Language Models
- NOVO: Bridging LLaVA and SAM with Visual-only Prompts for Reasoning Segmentation
- How Do VLAs Effectively Inherit from VLMs?
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- Rep2Text: Decoding Full Text from a Single LLM Token Representation
- Interaction-Centric Knowledge Infusion and Transfer for Open-Vocabulary Scene Graph Generation
- MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition
- VLDrive: Vision-Augmented Lightweight MLLMs for Efficient Language-grounded Autonomous Driving
- S2LM: Towards Semantic Steganography via Large Language Models
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Medical Referring Image Segmentation via Next-Token Mask Prediction
- DeepEyesV2: Toward Agentic Multimodal Model
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- Embodiment Transfer Learning for Vision-Language-Action Models
- Seeing Straight: Document Orientation Detection for Efficient OCR
- An LLM-based Framework for Human-Swarm Teaming Cognition in Disaster Search and Rescue
- Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
- Seeing What You Say: Expressive Image Generation from Speech
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
- TRACE: Textual Reasoning for Affordance Coordinate Extraction
- In-Context Adaptation of VLMs for Few-Shot Cell Detection in Optical Microscopy
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
- When One Modality Sabotages the Others: A Diagnostic Lens on Multimodal Reasoning
- UniChange: Unifying Change Detection with Multimodal Large Language Model
- OmniVLA: Physically-Grounded Multimodal VLA with Unified Multi-Sensor Perception for Robotic Manipulation
- GeoToken: Hierarchical Geolocalization of Images via Next Token Prediction
- VesSAM: Efficient Multi-Prompting for Segmenting Complex Vessel
- Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
- GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks
- Leveraging Hierarchical Image-Text Misalignment for Universal Fake Image Detection
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Text-guided Fine-Grained Video Anomaly Understanding
- LGCA: Enhancing Semantic Representation via Progressive Expansion
- FedReplay: A Feature Replay Assisted Federated Transfer Learning Framework for Efficient and Privacy-Preserving Smart Agriculture
- CompAgent: An Agentic Framework for Visual Compliance Verification
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- ODP-Bench: Benchmarking Out-of-Distribution Performance Prediction
- Languages are Modalities: Cross-Lingual Alignment via Encoder Injection
- Generating Accurate and Detailed Captions for High-Resolution Images
- NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- CATCH: A Modular Cross-domain Adaptive Template with Hook
- Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
- ConceptScope: Characterizing Dataset Bias via Disentangled Visual Concepts
- Dynamic VLM-Guided Negative Prompting for Diffusion Models
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- Instruction-based image editing: a survey on data, models, evaluation, and applications
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- Rad-JEPA 3D: Radiology Joint-Embedding Predictive Model for 3D Computed Tomography
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- Mitigating Modal Imbalance in Multimodal Reasoning
- Prior Directions: Why GUI Grounding Gets Locked in the Past
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation
- RoCA: Robust Cross-Domain End-to-End Autonomous Driving
- Compressed Video Aggregator: Content-driven Module for Efficient Micro-Video Recommendation
- Facial-Expression-Aware Prompting for Empathetic LLM Tutoring
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- RL makes MLLMs see better than SFT
- PRISMM-Bench: A Benchmark of Peer-Review Grounded Multimodal Inconsistencies
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Large Language Model for Verilog Code Generation: Literature Review and the Road Ahead
- EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- FruitProm: Probabilistic Maturity Estimation and Detection of Fruits and Vegetables
- OpenLVLM-MIA: A Controlled Benchmark Revealing the Limits of Membership Inference Attacks on Large Vision-Language Models
- QSVD: Efficient Low-rank Approximation for Unified Query-Key-Value Weight Compression in Low-Precision Vision-Language Models
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Kernelized Sparse Fine-Tuning with Bi-level Parameter Competition for Vision Models
- SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
- Invoice Information Extraction: Methods and Performance Evaluation
- RAVEN: Robust Advertisement Video Violation Temporal Grounding via Reinforcement Reasoning
- Modeling Expert Interactions in Sparse Mixture of Experts via Graph Structures
- Decentralized Multi-Agent Goal Assignment for Path Planning using Large Language Models
- Explainable Detection of AI-Generated Images with Artifact Localization Using Faster-Than-Lies and Vision-Language Models for Edge Devices
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- A Survey on Efficient Vision-Language-Action Models
- Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- Large language model-based task planning for service robots: A review
- Positional Preservation Embedding for Multimodal Large Language Models
- LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- Windsock is Dancing: Adaptive Multimodal Retrieval-Augmented Generation
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- SRSR: Enhancing Semantic Accuracy in Real-World Image Super-Resolution with Spatially Re-Focused Text-Conditioning
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- HARMONY: Hidden Activation Representations and Model Output-Aware Uncertainty Estimation for Vision-Language Models
- Mitigating Coordinate Prediction Bias from Positional Encoding Failures
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Integrating Genomics into Multimodal EHR Foundation Models
- Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- FineRS: Fine-grained Reasoning and Segmentation of Small Objects with Reinforcement Learning
- REMONI: An Autonomous System Integrating Wearables and Multimodal Large Language Models for Enhanced Remote Health Monitoring
- Vision Language Models for Dynamic Human Activity Recognition in Healthcare Settings
- VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set
- Enhanced MLLM Black-Box Jailbreaking Attacks and Defenses
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
- PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- LLM-Integrated Bayesian State Space Models for Multimodal Time-Series Forecasting
- ARGenSeg: Image Segmentation with Autoregressive Image Generation Model
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
- StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback
- Fake-in-Facext: Towards Fine-Grained Explainable DeepFake Analysis
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- From Masks to Worlds: A Hitchhiker's Guide to World Models
- Better Tokens for Better 3D: Advancing Vision-Language Modeling in 3D Medical Imaging
- GenColorBench: A Color Evaluation Benchmark for Text-to-Image Generation Models
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- CoSense-LLM: Semantics at the Edge with Cost- and Uncertainty-Aware Cloud-Edge Cooperation
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- Modal Aphasia: Can Unified Multimodal Models Describe Images From Memory?
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
- PruneHal: Reducing Hallucinations in Multi-modal Large Language Models through Adaptive KV Cache Pruning
- Tibetan Language and AI: A Comprehensive Survey of Resources, Methods and Challenges
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- Beyond Single Models: Mitigating Multimodal Hallucinations via Adaptive Token Ensemble Decoding
- Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- SafeCoop: Unravelling Full Stack Safety in Agentic Collaborative Driving
- Accelerating Vision Transformers with Adaptive Patch Sizes
- Unbiased Gradient Low-Rank Projection
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
- Multilingual Text-to-Image Person Retrieval via Bidirectional Relation Reasoning and Aligning
- Token-Level Inference-Time Alignment for Vision-Language Models
- Multimodal Safety Is Asymmetric: Cross-Modal Exploits Unlock Black-Box MLLMs Jailbreaks
- FineVision: Open Data Is All You Need
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- AION-1: Omnimodal Foundation Model for Astronomical Sciences
- MERIT: Modular Framework for Multimodal Misinformation Detection with Web-Grounded Reasoning
- Xiaoice: Training-Free Video Understanding via Self-Supervised Spatio-Temporal Clustering of Semantic Features
- From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
- Efficient Vision-Language-Action Models for Embodied Manipulation: A Systematic Survey
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- AUGUSTUS: An LLM-Driven Multimodal Agent System with Contextualized User Memory
- Train a Unified Multimodal Data Quality Classifier with Synthetic Data
- Comprehensive language-image pre-training for 3D medical image understanding
- Composition-Grounded Data Synthesis for Visual Reasoning
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Learning an Image Editing Model without Image Editing Pairs
- ChangingGrounding: 3D Visual Grounding in Changing Scenes
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
- Benchmarking Multimodal Large Language Models for Face Recognition
- Backdoor Unlearning by Linear Task Decomposition
- Cross-Scenario Unified Modeling of User Interests at Billion Scale
- xLLM Technical Report
- GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- Efficient Video Sampling: Pruning Temporally Redundant Tokens for Faster VLM Inference
- Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback
- Exploring Cross-Modal Flows for Few-Shot Learning
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- IAD-GPT: Advancing Visual Knowledge in Multimodal Large Language Model for Industrial Anomaly Detection
- Sequential Comics for Jailbreaking Multimodal Large Language Models via Structured Visual Storytelling
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Efficient Few-Shot Learning in Remote Sensing: Fusing Vision and Vision-Language Models
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- InfraGPT Smart Infrastructure: An End-to-End VLM-Based Framework for Detecting and Managing Urban Defects
- End-to-End Multi-Modal Diffusion Mamba
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- OralGPT: A Two-Stage Vision-Language Model for Oral Mucosal Disease Diagnosis and Description
- Rethinking the Simulation vs. Rendering Dichotomy: No Free Lunch in Spatial World Modelling
- Only-Style: Stylistic Consistency in Image Generation without Content Leakage
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Model-agnostic Adversarial Attack and Defense for Vision-Language-Action Models
- The Mechanistic Emergence of Symbol Grounding in Language Models
- Generative Universal Verifier as Multimodal Meta-Reasoner
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Vision Language Models Map Logos to Text via Semantic Entanglement in the Visual Projector
- Reflection-Based Task Adaptation for Self-Improving VLA
- Towards Cross-Modal Error Detection with Tables and Images
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- Evolution of meta's llama models and parameter-efficient fine-tuning of large language models: a survey
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- ViCO: A Training Strategy towards Semantic Aware Dynamic High-Resolution
- VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- UALM: Unified Audio Language Model for Understanding, Generation and Reasoning
- Evaluating Open-Source Vision-Language Models for Multimodal Sarcasm Detection
- Scaling Language-Centric Omnimodal Representation Learning
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- REACT3D: Recovering Articulations for Interactive Physical 3D Scenes
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- LSVOS 2025 Challenge Report: Recent Advances in Complex Video Object Segmentation
- CoDefend: Cross-Modal Collaborative Defense via Diffusion Purification and Prompt Optimization
- Enhancing Zero-Shot Anomaly Detection: CLIP-SAM Collaboration with Cascaded Prompts
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- A Survey on Agentic Multimodal Large Language Models
- Topological Alignment of Shared Vision-Language Embedding Space
- Prompt-Guided Spatial Understanding with RGB-D Transformers for Fine-Grained Object Relation Reasoning
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- Large Language Model-Empowered Channel Prediction and Predictive Beamforming for LEO Satellite Communications
- Unified Open-World Segmentation with Multi-Modal Prompts
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- Complementary and Contrastive Learning for Audio-Visual Segmentation
- Efficient Onboard Vision-Language Inference in UAV-Enabled Low-Altitude Economy Networks via LLM-Enhanced Optimization
- MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
- RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Tapered Language Models
- Nav-EE: Navigation-Guided Early Exiting for Efficient Vision-Language Models in Autonomous Driving
- CoIDO: Efficient Data Selection for Visual Instruction Tuning via Coupled Importance-Diversity Optimization
- Holistic Order Prediction in Natural Scenes
- MedQ-Bench: Evaluating and Exploring Medical Image Quality Assessment Abilities in MLLMs
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- Interpretable Generative and Discriminative Learning for Multimodal and Incomplete Clinical Data
- Entropy Meets Importance: A Unified Head Importance-Entropy Score for Stable and Efficient Transformer Pruning
- Zero-shot image privacy classification with Vision-Language Models
- RO-Bench: Large-scale robustness evaluation of MLLMs with text-driven counterfactual videos
- Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- Provable Training Data Identification for Large Language Models
- LM Fight Arena: Benchmarking Large Multimodal Models via Game Competition
- HandEval: Taking the First Step Towards Hand Quality Evaluation in Generated Images
- VisuoAlign: Safety Alignment of LVLMs with Multimodal Tree Search
- BLINK-Twice: You see, but do you observe? A Reasoning Benchmark on Visual Perception
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- NovaFlow: Zero-Shot Manipulation via Actionable Flow from Generated Videos
- MATRIX: Multimodal Agent Tuning for Robust Tool-Use Reasoning
- How to Teach Large Multimodal Models New Skills
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- LTCA: Long-range Temporal Context Attention for Referring Video Object Segmentation
- UniMMVSR: A Unified Multi-Modal Framework for Cascaded Video Super-Resolution
- MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- Multimodal Safety Evaluation in Generative Agent Social Simulations
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- VisualDAN: Exposing Vulnerabilities in VLMs with Visual-Driven DAN Commands
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- VideoCanvas: Unified Video Completion from Arbitrary Spatiotemporal Patches via In-Context Conditioning
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- MLLM4TS: Leveraging Vision and Multimodal Language Models for General Time-Series Analysis
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- ImageNet-Think-250K: A Large-Scale Synthetic Dataset for Multimodal Reasoning for Vision Language Models
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficiency in Audio LLMs
- SpotDiff: Spotting and Disentangling Interference in Feature Space for Subject-Preserving Image Generation
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- Detection and Measurement of Hailstones with Multimodal Large Language Models
- Gradient-Sign Masking for Task Vector Transport Across Pre-Trained Models
- VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- Syn-Diag: An LLM-based Synergistic Framework for Generalizable Few-shot Fault Diagnosis on the Edge
- SD-MVSum: Script-Driven Multimodal Video Summarization Method and Datasets
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Bridging Reasoning to Learning: Unmasking Illusions using Complexity Out of Distribution Generalization
- Beyond the Seen: Bounded Distribution Estimation for Open-Vocabulary Learning
- ContextNav: Towards Agentic Multimodal In-Context Learning
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- Contrastive Representation Regularization for Vision-Language-Action Models
- Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- Zephyrus: An Agentic Framework for Weather Science
- AgentTypo: Adaptive Typographic Prompt Injection Attacks against Black-box Multimodal Agents
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- SITCOM: Scaling Inference-Time COMpute for VLAs
- RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- UGround: Towards Unified Visual Grounding with Unrolled Transformers
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Efficient Test-Time Scaling for Small Vision-Language Models
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- Multimodal Carotid Risk Stratification with Large Vision-Language Models: Benchmarking, Fine-Tuning, and Clinical Insights
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- Brain-Language Model Alignment: Insights into the Platonic Hypothesis and Intermediate-Layer Advantage
- Backdoor Attacks Against Speech Language Models
- TextCAM: Explaining Class Activation Map with Text
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- TsLLM: Augmenting LLMs for General Time Series Understanding and Prediction
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- Data Selection for Fine-tuning Vision Language Models via Cross Modal Alignment Trajectories
- BioVERSE: Representation Alignment of Biomedical Modalities to LLMs for Multi-Modal Reasoning
- Glaucoma Detection and Structured OCT Report Generation via a Fine-tuned Multimodal Large Language Model
- Apriel-1.5-15b-Thinker
- Retrieval-Augmented Generation for Electrocardiogram-Language Models
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
- Catalog-Native LLM: Speaking Item-ID Dialect with Less Entanglement for Recommendation
- EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
- ProfVLM: A Lightweight Video-Language Model for Multi-View Proficiency Estimation
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- PatchVSR: Breaking Video Diffusion Resolution Limits with Patch-wise Video Super-Resolution
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- NuRisk: A Visual Question Answering Dataset for Agent-Level Risk Assessment in Autonomous Driving
- Understanding the Mixture-of-Experts with Nadaraya-Watson Kernel
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- dVLA: Diffusion Vision-Language-Action Model with Multimodal Chain-of-Thought
- DescribeEarth: Describe Anything for Remote Sensing Images
- EfficientVLA: Training-Free Acceleration and Compression for Vision-Language-Action Models
- Effective Model Pruning
- Towards Reliable and Holistic Visual In-Context Learning Prompt Selection
- PRPO: Paragraph-level Policy Optimization for Vision-Language Deepfake Detection
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- Judging by Appearances? Auditing and Intervening Vision-Language Models for Bail Prediction
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Defeating Cerberus: Concept-Guided Privacy-Leakage Mitigation in Multimodal Language Models
- Saliency Guided Longitudinal Medical Visual Question Answering
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- LayerD: Decomposing Raster Graphic Designs into Layers
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- From Code to Action: Hierarchical Learning of Diffusion-VLM Policies
- ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
- Performance-Efficiency Trade-off for Fashion Image Retrieval
- Uni-X: Mitigating Modality Conflict with a Two-End-Separated Architecture for Unified Multimodal Models
- When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- Latent Visual Reasoning
- World-Env: Leveraging World Model as a Virtual Environment for VLA Post-Training
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- DepthLM: Metric Depth From Vision Language Models
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- Mitigating Visual Hallucinations via Semantic Curriculum Preference Optimization in MLLMs
- FreeRet: MLLMs as Training-Free Retrievers
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
- Towards Redundancy Reduction in Diffusion Models for Efficient Video Super-Resolution
- Vision-Grounded Machine Interpreting: Improving the Translation Process through Visual Cues
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- AutoPrune: Each Complexity Deserves a Pruning Policy
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- PreScope: Unleashing the Power of Prefetching for Resource-Constrained MoE Inference
- VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis
- HunyuanImage 3.0 Technical Report
- Assessing Visual Privacy Risks in Multimodal AI: A Novel Taxonomy-Grounded Evaluation of Vision-Language Models
- Uncovering Grounding IDs: How External Cues Shape Multimodal Binding
- RIV: Recursive Introspection Mask Diffusion Vision Language Model
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- LOTFormer: Doubly-Stochastic Linear Attention via Low-Rank Optimal Transport
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- Tree Reward-Aligned Search for TReASURe in Masked Diffusion Language Models
- Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- Uncovering Intrinsic Capabilities: A Paradigm for Data Curation in Vision-Language Models
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Patient-specific Biomolecular Instruction Tuning
- MMPB: It's Time for Multi-Modal Personalization
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
- UniMIC: Token-Based Multimodal Interactive Coding for Human-AI Collaboration
- Explaining multimodal LLMs via intra-modal token interactions
- RAU: Reference-based Anatomical Understanding with Vision Language Models
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- Benchmarking and Mitigating Sycophancy in Medical Vision Language Models
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- DynaNav: Dynamic Feature and Layer Selection for Efficient Visual Navigation
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Discrete Guidance Matching: Exact Guidance for Discrete Flow Matching
- Abductive Logical Rule Induction by Bridging Inductive Logic Programming and Multimodal Large Language Models
- Training-Free Multimodal Deepfake Detection via Graph Reasoning
- UniVid: Unifying Vision Tasks with Pre-trained Video Generation Models
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- VLCE: A Knowledge-Enhanced Framework for Image Description in Disaster Assessment
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- Semantic Edge-Cloud Communication for Real-Time Urban Traffic Surveillance with ViT and LLMs over Mobile Networks
- TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- FORCE: Transferable Visual Jailbreaking Attacks via Feature Over-Reliance CorrEction
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- Null-Space Filtering for Data-Free Continual Model Merging: Preserving Transparency, Promoting Fidelity
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- ImaginationPolicy: Towards Generalizable, Precise and Reliable End-to-End Policy for Robotic Manipulation
- EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- JaiLIP: Jailbreaking Vision-Language Models via Loss Guided Image Perturbation
- EchoBench: Benchmarking Sycophancy in Medical Large Vision-Language Models
- Embodied AI: From LLMs to World Models
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
- Logics-Parsing Technical Report
- CAMILA: Context-Aware Masking for Image Editing with Language Alignment
- RSAVQ: Riemannian Sensitivity-Aware Vector Quantization for Large Language Models
- SIM-CoT: Supervised Implicit Chain-of-Thought
- Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
- Unifying Adversarially Robust Model Experts in Vision-Language Models
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language Models
- Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation
- LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
- ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- Low-bit Model Quantization for Deep Neural Networks: A Survey
- QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding
- Energy-Driven Adaptive Visual Token Pruning for Efficient Vision-Language Models
- Harnessing Synthetic Data from Generative AI for Statistical Inference
- Propose and Rectify: A Forensics-Driven MLLM Framework for Image Manipulation Localization
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- Alternating Training-based Label Smoothing Enhances Prompt Generalization
- PoRe: Position-Reweighted Visual Token Pruning for Vision Language Models
- The Platonic Universe: Do Foundation Models See the Same Sky?
- From Global to Local: Social Bias Transfer in CLIP
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- MAPO: Mixed Advantage Policy Optimization
- COLT: Enhancing Video Large Language Models with Continual Tool Usage
- Knowledge Transfer from Interaction Learning
- MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
- Steering Multimodal Large Language Models Decoding for Context-Aware Safety
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- GRPO++: Enhancing Dermatological Reasoning under Low Resource Settings
- The 1st Solution for MOSEv2 Challenge 2025: Long-term and Concept-aware Video Segmentation via SeC
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- ColorBlindnessEval: Can Vision-Language Models Pass Color Blindness Tests?
- Actions Speak Louder than Prompts: A Large-Scale Study of LLMs for Graph Inference
- Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
- Do Modern Video-LLMs Need to Listen? A Benchmark Audit and Scalable Remedy
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- V2V-GoT: Vehicle-to-Vehicle Cooperative Autonomous Driving with Multimodal Large Language Models and Graph-of-Thoughts
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Training-Free Label Space Alignment for Universal Domain Adaptation
- LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning about Source Code
- UIPro: Unleashing Superior Interaction Capability For GUI Agents
- Qwen3-Omni Technical Report
- Virtual Consistency for Audio Editing
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- Random Direct Preference Optimization for Radiography Report Generation
- MCTS-EP: Empowering Embodied Planning with Online Preference Optimization
- FitPro: A Zero-Shot Framework for Interactive Text-based Pedestrian Retrieval in Open World
- Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
- Causality-Induced Positional Encoding for Transformer-Based Representation Learning of Non-Sequential Features
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- A Vision-Language-Action-Critic Model for Robotic Real-World Reinforcement Learning
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- Pyramid Token Pruning for High-Resolution Large Vision-Language Models via Region, Token, and Instruction-Guided Importance
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
- EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery
- Robust Object Detection for Autonomous Driving via Curriculum-Guided Group Relative Policy Optimization
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Speech Language Models for Under-Represented Languages: Insights from Wolof
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- OmniMRI: A Unified Vision--Language Foundation Model for Generalist MRI Interpretation
- Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
- VLA-LPAF: Lightweight Perspective-Adaptive Fusion for Vision-Language-Action to Enable More Unconstrained Robotic Manipulation
- Towards Human-like Multimodal Conversational Agent by Generating Engaging Speech
- VisMoDAl: Visual Analytics for Evaluating and Improving Corruption Robustness of Vision-Language Models
- Unleashing the Potential of Multimodal LLMs for Zero-Shot Spatio-Temporal Video Grounding
- A Framework for Generating Artificial Datasets to Validate Absolute and Relative Position Concepts
- M-PACE: Mother Child Framework for Multimodal Compliance
- CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
- EDITS: Enhancing Dataset Distillation with Implicit Textual Semantics
- MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
- Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- GestOS: Advanced Hand Gesture Interpretation via Large Language Models to control Any Type of Robot
- Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Towards Rationale-Answer Alignment of LVLMs via Self-Rationale Calibration
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- More performant and scalable: Rethinking contrastive vision-language pre-training of radiology in the LLM era
- WHU-STree: A Multi-modal Benchmark Dataset for Street Tree Inventory
- Defense-to-Attack: Bypassing Weak Defenses Enables Stronger Jailbreaks in Vision-Language Models
- AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
- Data Scaling Laws for Radiology Foundation Models
- Large Language Model-Based Automatic Formulation for Stochastic Optimization Models
- MaRVL-QA: A Benchmark for Mathematical Reasoning over Visual Landscapes
- Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- 3D Aware Region Prompted Vision Language Model
- Embodied Navigation Foundation Model
- Spec-LLaVA: Accelerating Vision-Language Models with Dynamic Tree-Based Speculative Decoding
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- DRAG: Data Reconstruction Attack using Guided Diffusion
- DUAL-VAD: Dual Benchmarks and Anomaly-Focused Sampling for Video Anomaly Detection
- SpecVLM: Fast Speculative Decoding in Vision-Language Models
- Pathological Truth Bias in Vision-Language Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- The System Description of CPS Team for Track on Driving with Language of CVPR 2024 Autonomous Grand Challenge
- Traffic-MLLM: Curiosity-Regularized Supervised Learning for Traffic Scenario Case-Based Reasoning
- OpenUrban3D: Annotation-Free Open-Vocabulary Semantic Segmentation of Large-Scale Urban Point Clouds
- Detecting Text Manipulation in Images using Vision Language Models
- Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
- Adaptive Token Merging for Efficient Transformer Semantic Communication at the Edge
- Towards Understanding Visual Grounding in Visual Language Models
- Synthetic Homes: A Multimodal Generative AI Pipeline for Residential Building Data Generation under Data Scarcity
- Improving Personalized Search with Regularized Low-Rank Parameter Updates
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Curriculum-Based Multi-Tier Semantic Exploration via Deep Reinforcement Learning
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- Discovering Divergent Representations between Text-to-Image Models
- Generalized User-Oriented Image Semantic Coding Empowered by Large Vision-Language Model
- Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
- PlantExpertVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science
- Prompt-Driven Image Analysis with Multimodal Generative AI: Detection, Segmentation, Inpainting, and Interpretation
- Retrieval-Augmented VLMs for Multimodal Melanoma Diagnosis
- MITS: A Large-Scale Multimodal Benchmark Dataset for Intelligent Traffic Surveillance
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Attribute-based Object Grounding and Robot Grasp Detection with Spatial Reasoning
- Visual Representation Alignment for Multimodal Large Language Models
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
- Fine-Tuning Vision-Language Models for Visual Navigation Assistance
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Reconstruction Alignment Improves Unified Multimodal Models
- Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
- LLaDA-VLA: Vision Language Diffusion Action Models
- Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
- Teaching AI Stepwise Diagnostic Reasoning with Report-Guided Chain-of-Thought Learning
- Harnessing Object Grounding for Time-Sensitive Video Understanding
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- D-HUMOR: Dark Humor Understanding via Multimodal Open-ended Reasoning -- A Benchmark Dataset and Method
- An Explainable Deep Neural Network with Frequency-Aware Channel and Spatial Refinement for Flood Prediction in Sustainable Cities
- Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models
- BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model
- PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
- SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation
- Semantic-guided LoRA Parameters Generation
- PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
- Dual-Domain Perspective on Degradation-Aware Fusion: A VLM-Guided Robust Infrared and Visible Image Fusion Framework
- Towards Open World Detection: A Survey
- AnomalyLMM: Bridging Generative Knowledge and Discriminative Retrieval for Text-Based Person Anomaly Search
- Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions?
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- A Foundation Model for Chest X-ray Interpretation with Grounded Reasoning via Online Reinforcement Learning
- Vehicle-to-Infrastructure Collaborative Spatial Perception via Multimodal Large Language Models
- Visible Yet Unreadable: A Systematic Blind Spot of Vision Language Models Across Writing Systems
- Sample-efficient Integration of New Modalities into Large Language Models
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- Mitigating Multimodal Hallucinations via Gradient-based Self-Reflection
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Structuring GUI Elements through Vision Language Models: Towards Action Space Generation
- Behavioral Fingerprinting of Large Language Models
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- AutoDrive-R2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- An Investigation of Visual Foundation Models Robustness
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- Analysing the Language of Neural Audio Codecs
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
- OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
- Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
- CARIS: A Context-Adaptable Robot Interface System for Personalized and Scalable Human-Robot Interaction
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- Discrete Prompt Tuning via Recursive Utilization of Black-box Multimodal Large Language Model for Personalized Visual Emotion Recognition
- Make me an Expert: Distilling from Generalist Black-Box Models into Specialized Models for Semantic Segmentation
- Two Causes, Not One: Rethinking Omission and Fabrication Hallucinations in MLLMs
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- Target-Oriented Single Domain Generalization
- DriveQA: Passing the Driving Knowledge Test
- VoCap: Video Object Captioning and Segmentation from Any Prompt
- Domain Generalization in-the-Wild: Disentangling Classification from Domain-Aware Representations
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- Generalizable Object Re-Identification via Visual In-Context Prompting
- MedFoundationHub: A Lightweight and Secure Toolkit for Deploying Medical Vision Language Foundation Models
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- PathMR: Multimodal Visual Reasoning for Interpretable Pathology Diagnosis
- GRAFT: GRaPH and Table Reasoning for Textual Alignment -- A Benchmark for Structured Instruction Following and Visual Reasoning
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- Scalable Object Detection in the Car Interior With Vision Foundation Models
- MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models
- Do MLLMs Really Understand the Charts?
- Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
- How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- ZeST: an LLM-based Zero-Shot Traversability Navigation for Unknown Environments
- Can we make NeRF-based visual localization privacy-preserving?
- Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
- Tailored Teaching with Balanced Difficulty: Elevating Reasoning in Multimodal Chain-of-Thought via Prompt Curriculum
- PRISM: Robust VLM Alignment with Principled Reasoning for Integrated Safety in Multimodality
- Revisiting associative recall in modern recurrent models
- CLARIFY: A Specialist-Generalist Framework for Accurate and Lightweight Dermatological Visual Question Answering
- DemoBias: An Empirical Study to Trace Demographic Biases in Vision Foundation Models
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- ArgusCogito: Chain-of-Thought for Cross-Modal Synergy and Omnidirectional Reasoning in Camouflaged Object Segmentation
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- Virtual Community: An Open World for Humans, Robots, and Society
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision Models
- RynnEC: Bringing MLLMs into Embodied World
- Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- The 9th AI City Challenge
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
- Contrastive Representations for Temporal Reasoning
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- Say It, See It: A Systematic Evaluation on Speech-Based 3D Content Generation Methods in Augmented Reality
- Region-Level Context-Aware Multimodal Understanding
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- OVG-HQ: Online Video Grounding with Hybrid-modal Queries
- Bongard-RWR+: Real-World Representations of Fine-Grained Concepts in Bongard Problems
- Language models align with brain regions that represent concepts across modalities
- ImagiDrive: A Unified Imagination-and-Planning Framework for Autonomous Driving
- Probing the Representational Power of Sparse Autoencoders in Vision Models
- Audio Flamingo Sound-CoT Technical Report: Improving Chain-of-Thought Reasoning in Sound Understanding
- UAV-VL-R1: Generalizing Vision-Language Models via Supervised Fine-Tuning and Multi-Stage GRPO for UAV Visual Reasoning
- Are Large Pre-trained Vision Language Models Effective Construction Safety Inspectors?
- A Dataset for Distilling Knowledge Priors from Literature for Therapeutic Design
- Agentic Design Review System
- AddressVLM: Cross-view Alignment Tuning for Image Address Localization using Large Vision-Language Models
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- MANGO: Multimodal Attention-based Normalizing Flow Approach to Fusion Learning
- Bridging Modality Gaps in e-Commerce Products via Vision-Language Alignment
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
- January Food Benchmark (JFB): A Public Benchmark Dataset and Evaluation Suite for Multimodal Food Analysis
- Taking the next step with generative artificial intelligence: The transformative role of multimodal large language models in science education
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- Vision Generalist Model: A Survey
- SHALE: A Scalable Benchmark for Fine-grained Hallucination Evaluation in LVLMs
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- VertexRegen: Mesh Generation with Continuous Level of Detail
- OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
- AMRG: Extend Vision Language Models for Automatic Mammography Report Generation
- Cowpox: Towards the Immunity of VLM-based Multi-Agent Systems
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- KFFocus: Highlighting Keyframes for Enhanced Video Understanding
- GeoVLA: Empowering 3D Representations in Vision-Language-Action Models
- Re:Verse -- Can Your VLM Read a Manga?
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- Selective Contrastive Learning for Weakly Supervised Affordance Grounding
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- ACD-CLIP: Decoupling Representation and Dynamic Fusion for Zero-Shot Anomaly Detection
- Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- CoT-Pose: Chain-of-Thought Reasoning for 3D Pose Generation from Abstract Prompts
- MolmoAct: Action Reasoning Models that can Reason in Space
- Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
- MCITlib: Multimodal Continual Instruction Tuning Library and Benchmark
- ForensicsSAM: Toward Robust and Unified Image Forgery Detection and Localization Resisting to Adversarial Attack
- Find Them All: Unveiling MLLMs for Versatile Person Re-identification
- eMotions: A Large-Scale Dataset and Audio-Visual Fusion Network for Emotion Analysis in Short-form Videos
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
- Multimodal learning with next-token prediction for large multimodal models
- Text Embedded Swin-UMamba for DeepLesion Segmentation
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- Text-guided Visual Prompt DINO for Generic Segmentation
- CountQA: How Well Do MLLMs Count in the Wild?
- AnimateScene: Camera-controllable Animation in Any Scene
- Automatic Semantic Alignment of Flow Pattern Representations for Exploration with Large Language Models
- Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
- User-Intent-Driven Semantic Communication via Adaptive Deep Understanding
- MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- MoMA: A Mixture-of-Multimodal-Agents Architecture for Enhancing Clinical Prediction Modelling
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
- VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
- SPEX: A Vision-Language Model for Land Cover Extraction on Spectral Remote Sensing Images
- A Survey on Video Temporal Grounding with Multimodal Large Language Model
- Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
- Latent Expression Generation for Referring Image Segmentation and Grounding
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- Enhancing Dialogue Annotation with Speaker Characteristics Leveraging a Frozen LLM
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- RoboTron-Sim: Improving Real-World Driving via Simulated Hard-Case
- Training-Free Multimodal Large Language Model Orchestration
- Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding
- FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models
- S2M3: Split-and-Share Multi-Modal Models for Distributed Multi-Task Inference on the Edge
- From eye to AI: studying rodent social behavior in the era of machine Learning
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- Controllable Hybrid Captioner for Improved Long-form Video Understanding
- Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
- Sustainability assessment using multimodal AI agents
- LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs
- GeoShield: Safeguarding Geolocation Privacy from Vision-Language Models via Adversarial Perturbations
- FedPromo: Federated Lightweight Proxy Models at the Edge Bring New Domains to Foundation Models
- Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- SceneLoom: Communicating Data with Scene Context
- Would you let a humanoid play storytelling with your child? A usability study on LLM-powered narrative Human-Robot Interaction
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- FedVLA: Federated Vision-Language-Action Learning with Dual Gating Mixture-of-Experts for Robotic Manipulation
- Free-MoRef: Instantly Multiplexing Context Perception Capabilities of Video-MLLMs within Single Inference
- VLM4D: Towards Spatiotemporal Awareness in Vision Language Models
- Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- Bench2ADVLM: A Closed-Loop Benchmark for Vision-language Models in Autonomous Driving
- InspectVLM: Unified in Theory, Unreliable in Practice
- MLP Memory: A Retriever-Pretrained Memory for Large Language Models
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- LLaDA-MedV: Exploring Large Language Diffusion Models for Biomedical Image Understanding
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks
- SGCap: Decoding Semantic Group for Zero-shot Video Captioning
- MeshLLM: Empowering Large Language Models to Progressively Understand and Generate 3D Mesh
- ForenX: Towards Explainable AI-Generated Image Detection with Multimodal Large Language Models
- Trans-Adapter: A Plug-and-Play Framework for Transparent Image Inpainting
- ROVI: A VLM-LLM Re-Captioned Dataset for Open-Vocabulary Instance-Grounded Text-to-Image Generation
- How LLMs are Shaping the Future of Virtual Reality
- Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- Contact-Aware Amodal Completion for Human-Object Interaction via Multi-Regional Inpainting
- Multimodal Referring Segmentation: A Survey
- Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models
- GeoExplorer: Active Geo-localization with Curiosity-Driven Exploration
- RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
- TriP-LLM: A Tri-Branch Patch-wise Large Language Model Framework for Time-Series Anomaly Detection
- Efficient Masked Attention Transformer for Few-Shot Classification and Segmentation
- ART: Adaptive Relation Tuning for Generalized Relation Prediction
- Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
- UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
- Bidirectional Likelihood Estimation with Multi-Modal Large Language Models for Text-Video Retrieval
- Unveiling Super Experts in Mixture-of-Experts Large Language Models
- Adversarial-Guided Diffusion for Multimodal LLM Attacks
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- Towards Omnimodal Expressions and Reasoning in Referring Audio-Visual Segmentation
- ScreenCoder: Advancing Visual-to-Code Generation for Front-End Automation via Modular Multimodal Agents
- GeoReg: Weight-Constrained Few-Shot Regression for Socio-Economic Estimation using LLM
- Text-Aware Image Restoration with Diffusion Models
- Doctor Sun: A Bilingual Multimodal Large Language Model for Biomedical AI
- Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
- Invisible Injections: Exploiting Vision-Language Models Through Steganographic Prompt Embedding
- Vision-Language Cross-Attention for Real-Time Autonomous Driving
- DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
- SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
- HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
- Compression Strategies for Efficient Multimodal LLMs in Medical Contexts
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
- Meta CLIP 2: A Worldwide Scaling Recipe
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- Learning Transferable Facial Emotion Representations from Large-Scale Semantically Rich Captions
- METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
- Advancing Compositional LLM Reasoning with Structured Task Relations in Interactive Multimodal Communications
- TransPrune: Token Transition Pruning for Efficient Large Vision-Language Model
- Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression
- Annotation-Free Human Sketch Quality Assessment
- AgroBench: Vision-Language Model Benchmark in Agriculture
- Customize Multi-modal RAI Guardrails with Precedent-based predictions
- GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
- RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
- NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- Trust the Model: Compact VLMs as In-Context Judges for Image-Text Data Quality
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- CONCAP: Seeing Beyond English with Concepts Retrieval-Augmented Captioning
- The Devil is in the EOS: Sequence Training for Detailed Image Captioning
- Position: Reasoning After Perception Means Reasoning Without Vision
- TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
- Object-centric Video Question Answering with Visual Grounding and Referring
- Closing the Modality Gap for Mixed Modality Search
- Seeing Beyond Frames: Zero-Shot Pedestrian Intention Prediction with Raw Temporal Video and Multimodal Cues
- Orbis: Overcoming Challenges of Long-Horizon Prediction in Driving World Models
- A Survey of Multimodal Hallucination Evaluation and Detection
- Towards Effective Human-in-the-Loop Assistive AI Agents
- LMM-Det: Make Large Multimodal Models Excel in Object Detection
- Resource Consumption Red-Teaming for Large Vision-Language Models
- Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction
- Captain Cinema: Towards Short Movie Generation
- DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis
- Adapting Large VLMs with Iterative and Manual Instructions for Generative Low-light Enhancement
- Emotion Recognition from Skeleton Data: A Comprehensive Survey
- VAT-KG: Knowledge-Intensive Multimodal Knowledge Graph Dataset for Retrieval-Augmented Generation
- Grounding Degradations in Natural Language for All-In-One Video Restoration
- LeAdQA: LLM-Driven Context-Aware Temporal Grounding for Video Question Answering
- Med-GRIM: Enhanced Zero-Shot Medical VQA using prompt-embedded Multimodal Graph RAG
- Descrip3D: Enhancing Large Language Model-based 3D Scene Understanding with Object-Level Text Descriptions
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- Docopilot: Improving Multimodal Models for Document-Level Understanding
- Generative Distribution Distillation
- IRGPT: Understanding Real-world Infrared Image with Bi-cross-modal Curriculum on Large-scale Benchmark
- Towards Urban Planing AI Agent in the Age of Agentic AI
- Efficient Whole Slide Pathology VQA via Token Compression
- VizGenie: Toward Self-Refining, Domain-Aware Workflows for Next-Generation Scientific Visualization
- Moodifier: MLLM-Enhanced Emotion-Driven Image Editing
- FedVLMBench: Benchmarking Federated Fine-Tuning of Vision-Language Models
- Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
- VLM-Guided Visual Place Recognition for Planet-Scale Geo-Localization
- HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMs
- RoadBench: A Vision-Language Foundation Model and Benchmark for Road Damage Understanding
- R-Stitch: Dynamic Trajectory Stitching for Efficient Reasoning
- A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
- DNT: a Deeply Normalized Transformer that can be trained by Momentum SGD
- InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
- Semi-off-Policy Reinforcement Learning for Vision-Language Slow-Thinking Reasoning
- Automatic Fine-grained Segmentation-assisted Report Generation
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- Advancing Visual Large Language Model for Multi-granular Versatile Perception
- Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models
- Argus: Leveraging Multiview Images for Improved 3-D Scene Understanding With Large Language Models
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- GLAD: Generalizable Tuning for Vision-Language Models
- DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
- Compact Vision Transformer by Reduction of Kernel Complexity
- LaViPlan : Language-Guided Visual Path Planning with RLVR
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
- Insights into a radiology-specialised multimodal large language model with sparse autoencoders
- Aligning Knowledge Graphs and Language Models for Factual Accuracy
- Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
- Mitigating Object Hallucinations via Sentence-Level Early Intervention
- AutoVDC: Automated Vision Data Cleaning Using Vision-Language Models
- ReAL-AD: Towards Human-Like Reasoning in End-to-End Autonomous Driving
- 3D-MoRe: Unified Modal-Contextual Reasoning for Embodied Question Answering
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- MindJourney: Test-Time Scaling with World Models for Spatial Reasoning
- Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
- Cross-Frame Representation Alignment for Fine-Tuning Video Diffusion Models
- Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Hybrid Reasoning for Perception, Explanation, and Autonomous Action in Manufacturing
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
- mSTEB: Massively Multilingual Evaluation of LLMs on Speech and Text Tasks
- PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly
- Teach Me Sign: Stepwise Prompting LLM for Sign Language Production
- NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- Bridging the Gap in Vision Language Models in Identifying Unsafe Concepts Across Modalities
- UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
- Graph World Model
- CoralVQA: A Large-Scale Visual Question Answering Dataset for Coral Reef Image Understanding
- Multiple Choice Learning of Low-Rank Adapters for Language Modeling
- DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
- FaceLLM: A Multimodal Large Language Model for Face Understanding
- Synthesizing Near-Boundary OOD Samples for Out-of-Distribution Detection
- A Training-Free, Task-Agnostic Framework for Enhancing MLLM Performance on High-Resolution Images
- Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
- ADAM: Autonomous Discovery and Annotation Model using LLMs for Context-Aware Annotations
- Generative Audio Language Modeling with Continuous-valued Tokens and Masked Next-Token Prediction
- Foundation Models in Medical Imaging: A Review and Outlook
- DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models
- Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning
- LLM-Guided Agentic Object Detection for Open-World Understanding
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
- LLaVA-c: Continual Improved Visual Instruction Tuning
- Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
- SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
- Uncertainty-Driven Expert Control: Enhancing the Reliability of Medical Vision-Language Models
- PoseLLM: Enhancing Language-Guided Human Pose Estimation with MLP Alignment
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels
- Multilingual Multimodal Software Developer for Code Generation
- Large Multi-modal Model Cartographic Map Comprehension for Textual Locality Georeferencing
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- Raptor: Scalable Train-Free Embeddings for 3D Medical Volumes Leveraging Pretrained 2D Foundation Models
- InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems
- Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models
- Multigranular Evaluation for Brain Visual Decoding
- Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
- SURPRISE3D: A Dataset for Spatial Understanding and Reasoning in Complex 3D Scenes
- Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- ConsNoTrainLoRA: Data-driven Weight Initialization of Low-rank Adapters using Constraints
- ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation
- DisenQ: Disentangling Q-Former for Activity-Biometrics
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- Outcome-Guided Distillation: A Teacher-Student Framework to Advance VLM Reasoning in Autonomous Driving
- Text-promptable Object Counting via Quantity Awareness Enhancement
- Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
- LangSplatV2: High-dimensional 3D Language Gaussian Splatting with 450+ FPS
- Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
- MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning
- Robust Multimodal Large Language Models Against Modality Conflict
- Is Diversity All You Need for Scalable Robotic Manipulation?
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
- Rethinking Layered Graphic Design Generation with a Top-Down Approach
- FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval
- VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
- Beyond Simple Edits: X-Planner for Complex Instruction-Based Image Editing
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Differential Attention for Multimodal Crisis Event Analysis
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- Vision-Language Models Can't See the Obvious
- A Visual Leap in CLIP Compositionality Reasoning through Generation of Counterfactual Sets
- MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
- VLM2Vec-V2: Advancing Multimodal Embedding for Videos, Images, and Visual Documents
- A Careful Examination of Large Behavior Models for Multitask Dexterous Manipulation
- Pre-Trained Policy Discriminators are General Reward Models
- Evaluation-Verification Reward for Consistent Multi-Reference Image Editing
- DreamVLA: A Vision-Language-Action Model Dreamed with Comprehensive World Knowledge
- CoT-Segmenter: Enhancing OOD Detection in Dense Road Scenes via Chain-of-Thought Reasoning
- Demystifying ChatGPT: How It Masters Genre Recognition
- VICI: VLM-Instructed Cross-view Image-localisation
- MolVision: Molecular Property Prediction with Vision Language Models
- NOVO: Unlearning-Compliant Vision Transformers
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- LACONIC: A 3D Layout Adapter for Controllable Image Creation
- Beyond Accuracy: Metrics that Uncover What Makes a 'Good' Visual Descriptor
- Bootstrapping Grounded Chain-of-Thought in Multimodal LLMs for Data-Efficient Model Adaptation
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
- AIGI-Holmes: Towards Explainable and Generalizable AI-Generated Image Detection via Multimodal Large Language Models
- Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
- Intelligent Histology for Tumor Neurosurgery
- SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
- Understanding Trade offs When Conditioning Synthetic Data
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
- MoIRA: Modular Instruction Routing Architecture for Multi-Task Robotics
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
- TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
- CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
- DiffusionLight-Turbo: Accelerated Light Probes for Free via Single-Pass Chrome Ball Inpainting
- VLAD: A VLM-Augmented Autonomous Driving Framework with Hierarchical Planning and Interpretable Decision Process
- Beyond Overcorrection: Evaluating Diversity in T2I Models with DivBench
- Choose What to Manipulate: Revealing Data Scaling Laws in Bounding-Box Guided Policies for Semantic Manipulation
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- ONLY: One-Layer Intervention Sufficiently Mitigates Hallucinations in Large Vision-Language Models
- GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
- CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
- Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
- Ascending the Infinite Ladder: Benchmarking Spatial Deformation Reasoning in Vision-Language Models
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- VISTA: Open-Vocabulary, Task-Relevant Robot Exploration with Online Semantic Gaussian Splatting
- From Answers to Rationales: Self-Aligning Multimodal Reasoning with Answer-Oriented Chain-of-Thought
- VOCAL: Visual Odometry via ContrAstive Learning
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- MotionGPT3: Human Motion as a Second Modality
- SceneSplat++: A Large Dataset and Comprehensive Benchmark for Language Gaussian Splatting
- Flash-VStream: Efficient Real-Time Understanding for Long Video Streams
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- A Clinically-Grounded Two-Stage Framework for Renal CT Report Generation
- Dataset Distillation via Vision-Language Category Prototype
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- A Survey on Vision-Language-Action Models for Autonomous Driving
- Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
- Pyramidal Patchification Flow for Visual Generation
- Dreamland: Controllable World Creation with Simulator and Generative Models
- CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems
- Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
- MTADiffusion: Mask Text Alignment Diffusion Model for Object Inpainting
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
- Token Activation Map to Visually Explain Multimodal LLMs
- Where, What, Why: Towards Explainable Driver Attention Prediction
- CSBrain: A Cross-scale Spatiotemporal Brain Foundation Model for EEG Decoding
- UrbanLLaVA: A Multi-modal Large Language Model for Urban Intelligence with Spatial Reasoning and Understanding
- Empowering Small VLMs to Think with Dynamic Memorization and Exploration
- DGE-YOLO: Dual-Branch Gathering and Attention for Accurate UAV Object Detection
- Towards Explainable Bilingual Multimodal Misinformation Detection and Localization
- CLIP-like Model as a Foundational Density Ratio Estimator
- Decoupled Seg Tokens Make Stronger Reasoning Video Segmenter and Grounder
- ETA: Efficiency through Thinking Ahead, A Dual Approach to Self-Driving with Large Models
- ReCo: Reminder Composition Mitigates Hallucinations in Vision-Language Models
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
- Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning
- Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
- LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
- Universal Retrieval for Multimodal Trajectory Modeling
- 4D-VLA: Spatiotemporal Vision-Language-Action Pretraining with Cross-Scene Calibration
- R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- UniCA: Unified Covariate Adaptation for Time Series Foundation Model
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency
- Score2Instruct: Scaling Up Video Quality-Centric Instructions via Automated Dimension Scoring
- OmniEval: A Benchmark for Evaluating Omni-modal Models with Visual, Auditory, and Textual Inputs
- EVA: Mixture-of-Experts Semantic Variant Alignment for Compositional Zero-Shot Learning
- LLaVA-Pose: Enhancing Human Pose and Action Understanding via Keypoint-Integrated Instruction Tuning
- Curing Semantic Drift: A Dynamic Approach to Grounding Generation in Large Vision-Language Models
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning
- Exploring the Design Space of 3D MLLMs for CT Report Generation
- SharpZO: Hybrid Sharpness-Aware Vision Language Model Prompt Tuning via Forward-Only Passes
- Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
- ThinkSound: Chain-of-Thought Reasoning in Multimodal Large Language Models for Audio Generation and Editing
- MMSearch-R1: Incentivizing LMMs to Search
- LiteVLM: A Low-Latency Vision-Language Model Inference Pipeline for Resource-Constrained Environments
- Video Perception Models for 3D Scene Synthesis
- TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
- SFNet: Fusion of Spatial and Frequency-Domain Features for Remote Sensing Image Forgery Detection
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Towards Efficient Exemplar Based Image Editing with Multimodal VLMs
- iLearnRobot: An Interactive Learning-Based Multi-Modal Robot with Continuous Improvement
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Diffuse Everything: Multimodal Diffusion Models on Arbitrary State Spaces
- ZeroVO: Visual Odometry with Minimal Assumptions
- ChordPrompt: Orchestrating Cross-Modal Prompt Synergy for Multi-Domain Incremental Learning in CLIP
- Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
- Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents
- Robotic Perception with a Large Tactile-Vision-Language Model for Physical Property Inference
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- ToSA: Token Merging with Spatial Awareness
- Synthetic Visual Genome
- V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis
- Capturing Fine-Grained Alignments Improves 3D Affordance Detection
- MSR-Align: Policy-Grounded Multimodal Alignment for Safety-Aware Reasoning in Vision-Language Models
- Commander-GPT: Dividing and Routing for Multimodal Sarcasm Detection
- ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- Generalizing vision-language models to novel domains: A comprehensive survey
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- Escaping the SpuriVerse: Can Large Vision-Language Models Generalize Beyond Seen Spurious Correlations?
- ARD-LoRA: Dynamic Rank Allocation for Parameter-Efficient Fine-Tuning of Foundation Models with Heterogeneous Adaptation Needs
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
- See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- PDF Retrieval Augmented Question Answering
- PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- Can Generated Images Serve as a Viable Modality for Text-Centric Multimodal Learning?
- VLA-OS: Structuring and Dissecting Planning Representations and Paradigms in Vision-Language-Action Models
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
- Fast ECoT: Efficient Embodied Chain-of-Thought via Thoughts Reuse
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Open World Scene Graph Generation using Vision Language Models
- Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models
- Visual-Instructed Degradation Diffusion for All-in-One Image Restoration
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
- Proxy-Embedding as an Adversarial Teacher: An Embedding-Guided Bidirectional Attack for Referring Expression Segmentation Models
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- Evaluating Multimodal Large Language Models on Educational Textbook Question Answering
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- Aligning Text, Images, and 3D Structure Token-by-Token
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- Moment Sampling in Video LLMs for Long-Form Video QA
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- Show-o2: Improved Native Unified Multimodal Models
- InfiniPot-V: Memory-Constrained KV Cache Compression for Streaming Video Understanding
- PathCoT: Chain-of-Thought Prompting for Zero-shot Pathology Visual Reasoning
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
- DreamLight: Towards Harmonious and Consistent Image Relighting
- Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
- Dense360: Dense Understanding from Omnidirectional Panoramas
- VisText-Mosquito: A Unified Multimodal Dataset for Visual Detection, Segmentation, and Textual Explanation on Mosquito Breeding Sites
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- GRAM: A Generative Foundation Reward Model for Reward Generalization
- Segmenting Visuals With Querying Words: Language Anchors For Semi-Supervised Image Segmentation
- CEED-VLA: Consistency Vision-Language-Action Model with Early-Exit Decoding
- Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model
- FreeQ-Graph: Free-form Querying with Semantic Consistent Scene Graph for 3D Scene Understanding
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Mixture of Cognitive Reasoners: Modular Reasoning with Brain-Like Specialization
- ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- AdaLRS: Loss-Guided Adaptive Learning Rate Search for Efficient Foundation Model Pretraining
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- Evolution of ReID: From Early Methods to LLM Integration
- Discrete Diffusion in Large Language and Multimodal Models: A Survey
- Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation
- Dynamic Context-oriented Decomposition for Task-aware Low-rank Adaptation with Less Forgetting and Faster Convergence
- Qwen vs. Gemma Integration with Whisper: A Comparative Study in Multilingual SpeechLLM Systems
- PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue
- ROSA: Harnessing Robot States for Vision-Language and Action Alignment
- FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
- DualEdit: Dual Editing for Knowledge Updating in Vision-Language Models
- Dual-Priv Pruning : Efficient Differential Private Fine-Tuning in Multimodal Large Language Models
- Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
- Dynamic Modality Scheduling for Multimodal Large Models via Confidence, Uncertainty, and Semantic Consistency
- Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations
- Adaptive Batch-Wise Sample Scheduling for Direct Preference Optimization
- UAVs Meet Agentic AI: A Multidomain Survey of Autonomous Aerial Intelligence and Agentic UAVs
- LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
- The Safety Reminder: A Soft Prompt to Reactivate Delayed Safety Awareness in Vision-Language Models
- Interpretable Text-Guided Image Clustering via Iterative Search
- How Visual Representations Map to Language Feature Space in Multimodal LLMs
- Quizzard@INOVA Challenge 2025 -- Track A: Plug-and-Play Technique in Interleaved Multi-Image Model
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
- Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
- Foundation Models in Autonomous Driving: A Survey on Scenario Generation and Scenario Analysis
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- On the Natural Robustness of Vision-Language Models Against Visual Perception Attacks in Autonomous Driving
- VGR: Visual Grounded Reasoning
- Learning Compact Vision Tokens for Efficient Large Multimodal Models
- EgoPrivacy: What Your First-Person Camera Says About You?
- Auditing Data Provenance in Real-world Text-to-Image Diffusion Models for Privacy and Copyright Protection
- Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation
- RationalVLA: A Rational Vision-Language-Action Model with Dual System
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- DART: Differentiable Dynamic Adaptive Region Tokenizer for Vision Foundation Models
- Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
- PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
- LLMs Are Not Yet Ready for Deepfake Image Detection
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- GeoCAD: Local Geometry-Controllable CAD Generation with Large Language Models
- VideoExplorer: Think With Videos For Agentic Long-Video Understanding
- Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
- Graph-MLLM: Harnessing Multimodal Large Language Models for Multimodal Graph Learning
- Gondola: Grounded Vision Language Planning for Generalizable Robotic Manipulation
- LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- Towards Multimodal Graph Large Language Model
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- Harnessing Vision-Language Models for Time Series Anomaly Detection
- Stepwise Decomposition and Dual-stream Focus: A Novel Approach for Training-free Camouflaged Object Segmentation
- Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
- Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
- HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
- RecGPT: A Foundation Model for Sequential Recommendation
- Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models
- Collaborative Multi-LoRA Experts with Achievement-based Multi-Tasks Loss for Unified Multimodal Information Extraction
- Pts3D-LLM: Studying the Impact of Token Structure for 3D Scene Understanding With Large Language Models
- STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
- Towards an Explainable Comparison and Alignment of Feature Embeddings
- CoMemo: LVLMs Need Image Context with Image Memory
- Token Transforming: A Unified and Training-Free Token Compression Framework for Vision Transformer Acceleration
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- In-Context Learning for Label-Efficient Cancer Image Classification in Oncology
- MCA-Bench: A Multimodal Benchmark for Evaluating CAPTCHA Robustness Against VLM-based Attacks
- Refer to Any Segmentation Mask Group With Vision-Language Prompts
- VideoMolmo: Spatio-Temporal Grounding Meets Pointing
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
- Locality Preserving Markovian Transition for Instance Retrieval
- GeomHair: Reconstruction of Hair Strands from Colorless 3D Scans
- X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
- From Handwriting to Feedback: Evaluating VLMs and LLMs for AI-Powered Assessment in Indonesian Classrooms
- HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- FLAM: Frame-Wise Language-Audio Modeling
- Track Any Anomalous Object: A Granular Video Anomaly Detection Pipeline
- Towards Vision-Language-Garment Models for Web Knowledge Garment Understanding and Generation
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
- SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI Agents
- Scaling Laws for Robust Comparison of Open Foundation Language-Vision Models and Datasets
- Neural Network Reprogrammability: A Unified Theme on Model Reprogramming, Prompt Tuning, and Prompt Instruction
- From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- Degradation-Aware Image Enhancement via Vision-Language Classification
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- Structured Labeling Enables Faster Vision-Language Models for End-to-End Autonomous Driving
- SRD: Reinforcement-Learned Semantic Perturbation for Backdoor Defense in VLMs
- MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- AuthGuard: Generalizable Deepfake Detection via Language Guidance
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
- ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
- DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
- HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
- R3-VQA: "Read the Room" by Video Social Reasoning
- Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- PRJ: Perception-Retrieval-Judgement for Generated Images
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- Multimodal Tabular Reasoning with Privileged Structured Information
- SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing
- Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
- Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
- PCEvolve: Private Contrastive Evolution for Synthetic Dataset Generation via Few-Shot Private Data and Generative APIs
- MiMo-VL Technical Report
- Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads
- PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
- Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- Towards Geometry Problem Solving in the Large Model Era: A Survey
- Robustness in Both Domains: CLIP Needs a Robust Text Encoder
- HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
- Are Large Language Models Good Temporal Graph Learners?
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Seeing the Arrow of Time in Large Multimodal Models
- Kernel-based Unsupervised Embedding Alignment for Enhanced Visual Representation in Vision-language Models
- FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens
- InterRVOS: Interaction-aware Referring Video Object Segmentation
- Auto-Annotation with Expert-Crafted Guidelines: A Study through 3D LiDAR Detection Benchmark
- DPO Learning with LLMs-Judge Signal for Computer Use Agents
- Bridging Weakly-Supervised Learning and VLM Distillation: Noisy Partial Label Learning for Efficient Downstream Adaptation
- Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
- MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
- From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models
- Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos
- Dual-Process Image Generation
- IMAGHarmony: Controllable Image Editing with Consistent Object Quantity and Layout
- RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
- Cycle Consistency as Reward: Learning Image-Text Alignment without Human Preferences
- CogniAlign: Word-Level Multimodal Speech Alignment with Gated Cross-Attention for Alzheimer's Detection
- Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
- VRD-IU: Lessons from Visually Rich Document Intelligence and Understanding
- MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
- Large Language Models for EEG: A Comprehensive Survey and Taxonomy
- Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- 3DRS: MLLMs Need 3D-Aware Representation Supervision for Scene Understanding
- MoCA: Multi-modal Cross-masked Autoencoder for Time Series in Digital Health
- GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- Is Extending Modality The Right Path Towards Omni-Modality?
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- STORM: Benchmarking Visual Rating of MLLMs with a Comprehensive Ordinal Regression Dataset
- EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
- Align is not Enough: Multimodal Universal Jailbreak Attack against Multimodal Large Language Models
- OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
- Overcoming Multi-step Complexity in Multimodal Theory-of-Mind Reasoning: A Scalable Bayesian Planner
- PromptVFX: Text-Driven Fields for Open-World 3D Gaussian Animation
- NavBench: Probing Multimodal Large Language Models for Embodied Navigation
- PCoreSet: Effective Active Learning through Knowledge Distillation from Vision-Language Models
- ModuLM: Enabling Modular and Multimodal Molecular Relational Learning with Large Language Models
- SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
- Unlabeled Data Improves Fine-Grained Image Zero-shot Classification with Multimodal LLMs
- Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
- AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
- A Vision-Language Model for Focal Liver Lesion Classification
- Reinforced Correlation Between Vision and Language for Precise Medical AI Assistant
- SynSHRP2: A Synthetic Multimodal Benchmark for Driving Safety-critical Events Derived from Real-world Driving Data
- Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
- VaVLM: Toward Efficient Edge-Cloud Video Analytics With Vision-Language Models
- Task-Related Token Compression in Multimodal Large Language Models from an Explainability Perspective
- HueManity: Probing Fine-Grained Visual Perception in MLLMs
- Mitigating Image Captioning Hallucinations in Vision-Language Models
- Your Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate
- Common Inpainted Objects In-N-Out of Context
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks
- Temac: Multi-Agent Collaboration for Automated Web GUI Testing
- VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making
- Open CaptchaWorld: A Comprehensive Web-based Platform for Testing and Benchmarking Multimodal LLM Agents
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- EXP-Bench: Can AI Conduct AI Research Experiments?
- DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
- Reducing Annotation Burden in Physical Activity Research Using Vision-Language Models
- BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
- Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model
- Grid-LOGAT: Grid Based Local and Global Area Transcription for Video Question Answering
- DisTime: Distribution-based Time Representation for Video Large Language Models
- Bootstrapping LLM Robustness for VLM Safety via Reducing the Pretraining Modality Gap
- S4-Driver: Scalable Self-Supervised Driving Multimodal Large Language Modelwith Spatio-Temporal Visual Representation
- A Mathematical Perspective On Contrastive Learning
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time
- From Hallucinations to Jailbreaks: Rethinking the Vulnerability of Large Foundation Models
- Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning
- Advancing Compositional Awareness in CLIP with Efficient Fine-Tuning
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
- MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning
- Time Blindness: Why Video-Language Models Can't See What Humans Can?
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- GenSpace: Benchmarking Spatially-Aware Image Generation
- SORCE: Small Object Retrieval in Complex Environments
- When Large Multimodal Models Confront Evolving Knowledge: Challenges and Explorations
- Weakly-Supervised Affordance Grounding Guided by Part-Level Semantic Priors
- DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
- Preemptive Hallucination Reduction: An Input-Level Approach for Multimodal Language Model
- VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
- ZeroSep: Separate Anything in Audio with Zero Training
- Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model
- VModA: An Effective Framework for Adaptive NSFW Image Moderation
- Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
- A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluation Methods
- VISLIX: An XAI Framework for Validating Vision Models with Slice Discovery and Analysis
- FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
- Multi-Sourced Compositional Generalization in Visual Question Answering
- Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
- MOVi: Training-free Text-conditioned Multi-Object Video Generation
- QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- Stairway to Success: An Online Floor-Aware Zero-Shot Object-Goal Navigation Framework via LLM-Driven Coarse-to-Fine Exploration
- EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- Using Knowledge Graphs to harvest datasets for efficient CLIP model training
- AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning
- TrackVLA: Embodied Visual Tracking in the Wild
- Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?
- Spoken question answering for visual queries
- UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
- DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
- VidText: Towards Comprehensive Evaluation for Video Text Understanding
- StressTest: Can YOUR Speech LM Handle the Stress?
- Zero-Shot Vision Encoder Grafting via LLM Surrogates
- Towards Comprehensive Scene Understanding: Integrating First and Third-Person Views for LVLMs
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Farm-LightSeek: An Edge-centric Multimodal Agricultural IoT Data Analytics Framework with Lightweight LLMs
- Seeing the Threat: Vulnerabilities in Vision-Language Models to Adversarial Attack
- Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
- From Large AI Models to Agentic AI: A Tutorial on Future Intelligent Communications
- Fostering Video Reasoning via Next-Event Prediction
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
- VIRAL: Vision-grounded Integration for Reward design And Learning
- Universal Visuo-Tactile Video Understanding for Embodied Interaction
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
- Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
- Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying
- A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- Fork-Merge Decoding: Enhancing Multimodal Understanding in Audio-Visual Large Language Models
- SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing
- MSEarth: A Multimodal Scientific Dataset and Benchmark for Phenomena Uncovering in Earth Science
- PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series
- Pretraining Language Models to Ponder in Continuous Space
- AutoReproduce: Automatic AI Experiment Reproduction with Paper Lineage
- TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- Adapting Foundation Vision-Language Models to Medical Diagnosis via Query-Driven Expert Bridging
- Understand, Think, and Answer: Advancing Visual Reasoning with Large Multimodal Models
- DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
- PARTONOMY: Large Multimodal Models with Part-Level Visual Understanding
- OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
- AVCD: Mitigating Hallucinations in Audio-Visual Large Language Models through Contrastive Decoding
- ConText-CIR: Learning from Concepts in Text for Composed Image Retrieval
- RefAV: Towards Planning-Centric Scenario Mining
- Adversarial Attacks against Closed-Source MLLMs via Feature Optimal Alignment
- QwT-v2: Practical, Effective and Efficient Post-Training Quantization
- Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
- VLM Can Be a Good Assistant: Enhancing Embodied Visual Tracking with Self-Improving Vision-Language Models
- HoPE: Hybrid of Position Embedding for Long Context Vision-Language Models
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
- MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects
- MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- Attention! Your Vision Language Model Could Be Maliciously Manipulated
- Modeling Beyond MOS: Quality Assessment Models Must Integrate Context, Reasoning, and Multimodality
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- LlamaSeg: Image Segmentation via Autoregressive Mask Generation
- Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
- Can Visual Encoder Learn to See Arrows?
- Causal-LLaVA: Causal Disentanglement for Mitigating Hallucination in Multimodal Large Language Models
- NEXT: Multi-Grained Mixture of Experts via Text-Modulation for Multi-Modal Object Re-Identification
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- Benign-to-Toxic Jailbreaking: Inducing Harmful Responses from Harmless Prompts
- Grounding Language with Vision: A Conditional Mutual Information Calibrated Decoding Strategy for Reducing Hallucinations in LVLMs
- VLMLight: Safety-Critical Traffic Signal Control via Vision-Language Meta-Control and Dual-Branch Reasoning Architecture
- RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
- Time Series Generation Under Data Scarcity: A Unified Generative Modeling Approach
- Multimodal Reasoning Agent for Zero-Shot Composed Image Retrieval
- Efficient Multi-modal Long Context Learning for Training-free Adaptation
- ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking
- LISAT: Language-Instructed Segmentation Assistant for Satellite Imagery
- SEFE: Superficial and Essential Forgetting Eliminator for Multimodal Continual Instruction Tuning
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion Recognition
- FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
- AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models
- EMAC+: Embodied Multimodal Agent for Collaborative Planning with VLM+LLM
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
- TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
- Modality Equilibrium Matters: Minor-Modality-Aware Adaptive Alternating for Cross-Modal Memory Enhancement
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Multiplicity is an Inevitable and Inherent Challenge in Multimodal Learning
- Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model
- AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation
- Medical Large Vision Language Models with Multi-Image Visual Ability
- VL-SAM-V2: Open-World Object Detection with General and Specific Query Fusion
- Increasing Computation Resolves Conflicts in Vision Language Models
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
- ReFineVLA: Reasoning-Aware Teacher-Guided Transfer Fine-Tuning
- Beyond Editing Pairs: Fine-Grained Instructional Image Editing via Multi-Scale Learnable Regions
- MIND-Edit: MLLM Insight-Driven Editing via Language-Vision Projection
- DREAM: Drafting with Refined Target Features and Entropy-Adaptive Cross-Attention Fusion for Multimodal Speculative Decoding
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
- Assessing the Benefits of Combining Advanced Deep Learning Techniques for Post-Disaster Building Damage Assessment from UAV Imagery
- POQD: Performance-Oriented Query Decomposer for Multi-vector retrieval
- Protein Design with Dynamic Protein Vocabulary
- Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
- SAMA: Towards Multi-Turn Referential Grounded Video Chat with Large Language Models
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- Affective Image Editing: Shaping Emotional Factors via Text Descriptions
- So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection
- VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
- Linear Multi-Timescale Retention as a Memory-Efficient Vision-Language Bridge
- BiomechGPT: Extending Motion-Language Models to Clinical Motion Understanding
- Rethinking Causal Mask Attention for Vision-Language Inference
- v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
- Localizing Knowledge in Diffusion Transformers
- Advertising in AI systems: Society must be vigilant
- Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
- DiffusionReward: Enhancing Blind Face Restoration through Reward Feedback Learning
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations
- R-Genie: Reasoning-Guided Generative Image Editing
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data
- SVL: Empowering Spiking Neural Networks for Efficient 3D Open-World Understanding
- Towards General Continuous Memory for Vision-Language Models
- HoloLLM: Multisensory Foundation Model for Language-Grounded Human Sensing and Reasoning
- Enhancing Large Vision-Language Models with Layout Modality for Table Question Answering on Japanese Annual Securities Reports
- CAS-IQA: Teaching Vision-Language Models for Synthetic Angiography Quality Assessment
- Integrating Visual Interpretation and Linguistic Reasoning for Math Problem Solving
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
- OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
- Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks
- REOBench: Benchmarking Robustness of Earth Observation Foundation Models
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- From Evaluation to Defense: Advancing Safety in Video Large Language Models
- SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
- Robustifying Vision-Language Models via Dynamic Token Reweighting
- Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- SOLVE: Synergy of Language-Vision and End-to-End Networks for Autonomous Driving
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- Training-Free Reasoning and Reflection in MLLMs
- Efficient Motion Prompt Learning for Robust Visual Tracking
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- Backdoor Cleaning without External Guidance in MLLM Fine-tuning
- ALTo: Adaptive-Length Tokenizer for Autoregressive Mask Generation
- When VLMs Meet Image Classification: Test Sets Renovation via Missing Label Identification
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- Panoptic Captioning: An Equivalence Bridge for Image and Text
- ManipLVM-R1: Reinforcement Learning for Reasoning in Embodied Manipulation with Large Vision-Language Models
- MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
- Beyond Needle(s) in the Embodied Haystack: Environment, Architecture, and Training Considerations for Long Context Reasoning
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- MMaDA: Multimodal Large Diffusion Language Models
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- Beyond Hard and Soft: Hybrid Context Compression for Balancing Local and Global Information Retention
- Exploring The Visual Feature Space for Multimodal Neural Decoding
- NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
- Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
- ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
- P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
- Flashback: Memory-Driven Zero-shot, Real-time Video Anomaly Detection
- DiffPrune: differentiable information throttling for token pruning in vision-language models
- Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
- InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
- Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
- Robo-DM: Data Management For Large Robot Datasets
- Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
- Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
- OViP: Online Vision-Language Preference Learning for VLM Hallucination
- FAU at ImageCLEF 2026 Task on Multimodal Reasoning Robust Candidate Scoring and Concise Multilingual Visual Answering
- Large Language Models Implicitly Learn to See and Hear Just By Reading
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept Tokens
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- CAD-Coder: An Open-Source Vision-Language Model for Computer-Aided Design Code Generation
- Vision-Language Modeling Meets Remote Sensing: Models, Datasets and Perspectives
- VoQA: Visual-only Question Answering
- Ground-V: Teaching VLMs to Ground Complex Instructions in Pixels
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- MonitorVLM-v2: A Deployed Vision-Language Framework for Real-Time Safety Violation Detection
- AudioJailbreak: Jailbreak Attacks against End-to-End Large Audio-Language Models
- Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
- UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning
- Visual Instruction Bottleneck Tuning
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Beyond Words: Multimodal LLM Knows When to Speak
- ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
- RA-Touch: Retrieval-Augmented Touch Understanding with Enriched Visual Data
- Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
- Sat2Sound: A Unified Framework for Zero-Shot Soundscape Mapping
- SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information
- Benchmarking and Confidence Evaluation of LALMs For Temporal Reasoning
- AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning
- FLASH: Latent-Aware Semi-Autoregressive Speculative Decoding for Multimodal Tasks
- BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
- Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
- Understanding Complexity in VideoQA via Visual Program Generation
- VLC Fusion: Vision-Language Conditioned Sensor Fusion for Robust Object Detection
- Large Language Models and Their Applications in Roadway Safety and Mobility Enhancement: A Comprehensive Review
- GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
- Scalable Video-to-Dataset Generation for Cross-Platform Mobile Agents
- From Local Details to Global Context: Advancing Vision-Language Models with Attention-Based Selection
- MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
- BadNAVer: Exploring Jailbreak Attacks On Vision-and-Language Navigation
- STAR: Stage-Wise Attention-Guided Token Reduction for Efficient Large Vision-Language Models Inference
- From Shots to Stories: LLM-Assisted Video Editing with Unified Language Representations
- LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding
- NeuroGen: Neural Network Parameter Generation via Large Language Models
- Spatial-LLaVA: Enhancing Large Language Models with Spatial Referring Expressions for Visual Understanding
- Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
- CompBench: Benchmarking Complex Instruction-guided Image Editing
- PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- PRS-Med: Position Reasoning Segmentation in Medical Imaging
- GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
- iSegMan: Interactive Segment-and-Manipulate 3D Gaussians
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Are vision language models robust to uncertain inputs?
- UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings
- FIGhost: Fluorescent Ink-based Stealthy and Flexible Backdoor Attacks on Physical Traffic Sign Recognition
- Extracting Explainable Dates From Medical Images By Reverse-Engineering UNIX Timestamps
- Dynam3D: Dynamic Layered 3D Tokens Empower VLM for Vision-and-Language Navigation
- Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild
- Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
- Efficient Orthogonal Fine-Tuning with Principal Subspace Adaptation
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- Where did the ambiguity go? Examining how multimodal models interpret polysemous words
- Assistant Placement Aria: A Benchmark for Egocentric Placement Assistance
- Image-Space Rule Discovery
- SurgXBench: Explainable Vision-Language Model Benchmark for Surgery
- Content Generation Models in Computational Pathology: A Comprehensive Survey on Methods, Applications, and Challenges
- Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
- Beyond Modality Collapse: Representations Blending for Multimodal Dataset Distillation
- GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
- EdgeMM: Multi-Core CPU with Heterogeneous AI-Extension and Activation-aware Weight Pruning for Multimodal LLMs at Edge
- ToDMA: Large Model-Driven Massive Token Communications for Semantic Multiple Access
- Patho-R1: A Multimodal Reinforcement Learning-Based Pathology Expert Reasoner
- A Light and Smart Wearable Platform with Multimodal Foundation Model for Enhanced Spatial Reasoning in People with Blindness and Low Vision
- SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
- HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
- Sage Deer: A Super-Aligned Driving Generalist Is Your Copilot
- MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
- ChronoSteer: Bridging Large Language Model and Time Series Foundation Model via Synthetic Data
- Multi-Token Prediction Needs Registers
- Does Feasibility Matter? Understanding the Impact of Feasibility on Synthetic Training Data
- Task-Core Memory Management and Consolidation for Long-term Continual Learning
- Learning to See Locally and Align Clinically with Pathology Semantics for Radiology Report Generation
- Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation
- Air-Ground Collaboration for Language-Specified Missions in Unknown Environments
- Zero-shot Quantization: A Comprehensive Survey
- ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
- Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models
- Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
- OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning
- From Seeing to Doing: Bridging Reasoning and Decision for Robotic Manipulation
- ORACLE-Grasp: Zero-Shot Task-Oriented Robotic Grasping using Large Multimodal Models
- An integrated language-vision foundation model for conversational diagnostics and triaging in primary eye care
- STORYANCHORS: Generating Consistent Multi-Scene Story Frames for Long-Form Narratives
- CoLT: Teaching Multi-Modal Models to Think with Chain of Latent Thoughts
- VLM-KG: Multimodal Radiology Knowledge Graph Generation
- CLTP: Contrastive Language-Tactile Pre-training for 3D Contact Geometry Understanding
- DSADF: Thinking Fast and Slow for Decision Making
- Mitigating Visual Hallucinations in Multimodal Systems through Retrieval-Augmented Reliability-Aware Inference
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- HorusEye: Language as Dynamic Attention for Emergency Visual Analysis
- Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
- What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
- Visually Interpretable Subtask Reasoning for Visual Question Answering
- SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing
- Simple yet Effective Semi-supervised Knowledge Distillation from Vision-Language Models via Dual-Head Optimization
- QuantX: A Framework for Hardware-Aware Quantization of Generative AI Workloads
- Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning
- EmoVLM-KD: Fusing Distilled Expertise with Vision-Language Models for Visual Emotion Analysis
- DexWild: Dexterous Human Interactions for In-the-Wild Robot Policies
- DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
- Hallucination-Aware Multimodal Benchmark for Gastrointestinal Image Analysis with Large Vision-Language Models
- Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
- Visual Instruction Tuning with Chain of Region-of-Interest
- DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
- Exploring Multimodal Foundation AI and Expert-in-the-Loop for Sustainable Management of Wild Salmon Fisheries in Indigenous Rivers
- ICU-Bench:Benchmarking Continual Unlearning in Multimodal Large Language Models
- The ML.ENERGY Benchmark: Toward Automated Inference Energy Measurement and Optimization
- UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
- PlaceIt3D: Language-Guided Object Placement in Real 3D Scenes
- FG-CLIP: Fine-Grained Visual and Textual Alignment
- Hierarchical Pre-Training of Vision Encoders with Large Language Model
- PADriver: Towards Personalized Autonomous Driving
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Adaptive Markup Language Generation for Contextually-Grounded Visual Document Understanding
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- DriveAgent: Multi-Agent Structured Reasoning with LLM and Multimodal Sensor Fusion for Autonomous Driving
- R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation
- Compositional Image-Text Matching and Retrieval by Grounding Entities
- Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
- UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving
- Enhancing the Learning Experience: Using Vision-Language Models to Generate Questions for Educational Videos
- RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
- TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action
- Improving Editability in Image Generation with Layer-wise Memory
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Robotic Visual Instruction
- Towards Autonomous Micromobility through Scalable Urban Simulation
- MINERVA: Evaluating Complex Video Reasoning
- ROSClaw: An OpenClaw ROS 2 Framework for Agentic Robot Control and Interaction
- Experimental study on surveillance video-based indoor occupancy measurement with occupant-centric control
- R-C2: Cycle-Consistent Reinforcement Learning Improves Multimodal Reasoning
- AgentRVOS: Reasoning over Object Tracks for Zero-Shot Referring Video Object Segmentation
- Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models
- Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality?
- Diverse Semantics-Guided Feature Alignment and Decoupling for Visible-Infrared Person Re-Identification
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- Vision as Unified Multimodal Generation
- From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
- JointDiT: Enhancing RGB-Depth Joint Modeling with Diffusion Transformers
- ScaleTrack: Scaling and back-tracking Automated GUI Agents
- Pushing the Limits of Low-Bit Optimizers: A Focus on EMA Dynamics
- ReMem-VLA: Empowering Vision-Language-Action Model with Memory via Dual-Level Recurrent Queries
- Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation
- MERLIN: Building Low-SNR Robust Multimodal LLMs for Electromagnetic Signals
- Lightweight Visual Reasoning for Socially-Aware Robots
- FireRed-OCR Technical Report
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- A data- and compute-efficient chest X-ray foundation model beyond aggressive scaling
- Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
- LEVANTE-bench: Multi-Scale Comparison of VLMs to Children Using Cognitive Tasks (or, "Is Your VLM Smarter Than a 5th Grader?")
- Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
- COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning
- The Design Space of Tri-Modal Masked Diffusion Models
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- Mcity Data Engine: Iterative Model Improvement Through Open-Vocabulary Data Selection
- UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation
- Zoomer: Adaptive Image Focus Optimization for Black-box MLLM
- Towards Film-Making Production Dialogue, Narration, Monologue Adaptive Moving Dubbing Benchmarks
- Iterative Tool Usage Exploration for Multimodal Agents via Step-wise Preference Tuning
- OpenAVS: Training-Free Open-Vocabulary Audio Visual Segmentation with Foundational Models
- Responsive DNN Adaptation for Video Analytics against Environment Shift via Hierarchical Mobile-Cloud Collaborations
- VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive Interaction
- Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs
- DEL: Digit Entropy Loss for Numerical Learning of Large Language Models
- MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery
- SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
- OmicsLM: A Multimodal Large Language Model for Multi-Sample Omics Reasoning
- A Dialogue-Based Framework for Correcting Multimodal Errors in AI-Assisted STEM Education
- X-Fusion: Introducing New Modality to Frozen Large Language Models
- Classifier-to-Bias: Toward Unsupervised Automatic Bias Detection for Visual Classifiers
- CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
- UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
- Antidote: A Unified Framework for Mitigating LVLM Hallucinations in Counterfactual Presupposition and Object Perception
- Multimodal Large Language Models for Medicine: A Comprehensive Survey
- YoChameleon: Personalized Vision and Language Generation
- CompleteMe: Reference-based Human Image Completion
- Learning Streaming Video Representation via Multitask Training
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- Learning Action Priors for Cross-embodiment Robot Manipulation
- Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling
- AREA: Attribute Extraction and Aggregation for CLIP-Based Class-Incremental Learning
- EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- VLA Foundry: A Unified Framework for Training Vision-Language-Action Models
- Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks
- MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image
- Let ViT Speak: Generative Language-Image Pre-training
- Describe-Then-Act: Proactive Agent Steering via Distilled Language-Action World Models
- From Kepler to Newton: Inductive Biases Guide Learned World Models in Transformers
- SPARC: Separating Perception And Reasoning Circuits for Test-time Scaling of VLMs
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- Generative Visual Code Mobile World Models
- Mind the Gap No More: Achieving Zero-Gap Multimodal Integration via One Tokenizer
- Zero-shot adaptable task planning for autonomous construction robots: a comparative study of lightweight single and multi-AI agent systems
- VCM: Vision Concept Modeling Based on Implicit Contrastive Learning with Vision-Language Instruction Fine-Tuning
- PhenoAssistant: A Conversational Multi-Agent AI System for Automated Plant Phenotyping
- LLM-Powered GUI Agents in Phone Automation: Surveying Progress and Prospects
- VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
- Temporal UI State Inconsistency in Desktop GUI Agents: Formalizing and Defending Against TOCTOU Attacks on Computer-Use Agents
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- KIRA: Knowledge-Intensive Image Retrieval and Reasoning Architecture for Specialized Visual Domains
- GTA-2: Benchmarking General Tool Agents from Atomic Tool-Use to Open-Ended Workflows
- DeepInsert: Early Layer Bypass for Efficient and Performant Multimodal Understanding
- RadioFormer: A Multiple-Granularity Radio Map Estimation Transformer with 1\textpertenthousand Spatial Sampling
- Generative AI in Embodied Systems: System-Level Analysis of Performance, Efficiency and Scalability
- Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
- Unsupervised Visual Chain-of-Thought Reasoning via Preference Optimization
- Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
- StarVLA-α: Reducing Complexity in Vision-Language-Action Systems
- Mosaic: Cross-Modal Clustering for Efficient Video Understanding
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- SMART: When is it Actually Worth Expanding a Speculative Tree?
- MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
- Reinforced Attention Learning
- Semantic Leakage from Image Embeddings
- Toward Fully Autonomous Driving: AI, Challenges, Opportunities, and Needs
- ToolTok: Tool Tokenization for Efficient and Generalizable GUI Agents
- Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- The Semantic Lifecycle in Embodied AI: Acquisition, Representation and Storage via Foundation Models
- Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training
- ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
- Geometric Cross-Modal Token Selection for Latency-Constrained Multimodal Token Communication
- NeuroMosaic: Anatomically Grounded Multimodal Large Language Modeling for Molecularly Aware Glioma Reasoning from 3D MRI and Clinical Narratives
- OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
- DRC: Enhancing Personalized Image Generation via Disentangled Representation Composition
- From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
- UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
- Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding
- T2VAttack: Adversarial Attack on Text-to-Video Diffusion Models
- Ultra Lowrate Image Compression with Semantic Residual Coding and Compression-aware Diffusion
- Dual Prompting Image Restoration with Diffusion Transformers
- Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
- Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles
- DAC-Pose: Dual-Agent Collaborative Framework for Pose-Guided Human Generation
- CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly
- Evaluating Graphical Perception with Multimodal LLMs
- URECA: Unique Region Caption Anything
- Ternarization of Vision Language Models for use on edge devices
- Multimodal Alignment Through Joint Kernel Entropic Gromov--Wasserstein Optimal Transport
- TriCLE: Tri-Modal Vision-Language Reasoning for Edge-Deployed Fine-Grained Clustering
- SEAR: Simple and Efficient Adaptation of Visual Geometric Transformers for Unpaired RGB+Thermal 3D Reconstruction
- Taxonomy-Aware Evaluation of Vision-Language Models
- SCAM: A Real-World Typographic Robustness Evaluation for Multimodal Foundation Models
- Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
- Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
- A Survey of Foundation Model-Powered Recommender Systems: From Feature-Based, Generative to Agentic Paradigms
- Describe Anything: Detailed Localized Image and Video Captioning
- MedM-VL: What Makes a Good Medical LVLM?
- Domain Generalization for Face Anti-spoofing via Content-aware Composite Prompt Engineering
- OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
- Vidi: Large Multimodal Models for Video Understanding and Editing
- FaceInsight: A Multimodal Large Language Model for Face Perception
- AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
- TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
- VISTA-OCR: Towards generative and interactive end to end OCR models
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- AffordanceSAM: Segment Anything Once More in Affordance Grounding
- Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- Multimodal Large Language Models for Enhanced Traffic Safety: A Comprehensive Review and Future Trends
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
- AGI Is Coming... Right After AI Learns to Play Wordle
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- M2IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
- UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
- Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
- Modality Selection and Skill Segmentation via Cross-Modality Attention
- Learning from Reasoning Failures via Synthetic Data Generation
- LGD: Leveraging Generative Descriptions for Zero-Shot Referring Image Segmentation
- ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task
- ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- Manipulating Multimodal Agents via Cross-Modal Prompt Injection
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- Scaling LLaNA: Advancing NeRF-Language Understanding Through Large-Scale Training
- Analysing the Robustness of Vision-Language-Models to Common Corruptions
- Designing a reliable lateral movement detector using a graph foundation model
- Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
- Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
- Harmony: A Unified Framework for Modality Incremental Learning
- VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Post-pre-training for Modality Alignment in Vision-Language Foundation Models
- VistaDPO: Video Hierarchical Spatial-Temporal Direct Preference Optimization for Large Video Models
- JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image Restoration
- Window Token Concatenation for Efficient Visual Large Language Models
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
- Efficient Contrastive Decoding with Probabilistic Hallucination Detection - Mitigating Hallucinations in Large Vision Language Models -
- TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
- Domain-Grounded Candidate Selection for Agentic Image Editing: A Shadow Removal Case
- Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval
- Media Meets Communication in 6G: Fundamentals, Key Technologies, and Applications
- HoloCount: A Holistic Visual Counting Benchmark for MLLMs
- TARAC: Mitigating Hallucination in LVLMs via Temporal Attention Real-time Accumulative Connection
- VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models
- CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
- Imagine How To Change: Explicit Procedure Modeling for Change Captioning
- SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
- PVUW 2025 Challenge Report: Advances in Pixel-level Understanding of Complex Videos in the Wild
- Video Summarization with Large Language Models
- MIEB: Massive Image Embedding Benchmark
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- LangPert: Detecting and Handling Task-level Perturbations for Robust Object Rearrangement
- Socratic Chart: Cooperating Multiple Agents for Robust SVG Chart Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
- Efficient Prompt Tuning for Hierarchical Ingredient Recognition
- Improving Multimodal Hateful Meme Detection Exploiting LMM-Generated Knowledge
- ReasonDrive: Efficient Visual Question Answering for Autonomous Vehicles with Reasoning-Enhanced Small Vision-Language Models
- Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
- VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
- AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- AirVista-II: An Agentic System for Embodied UAVs Toward Dynamic Scene Semantic Understanding
- GenEDA: Towards Generative Netlist Functional Reasoning via Cross-Modal Circuit Encoder-Decoder Alignment
- HM-RAG: Hierarchical Multi-Agent Multimodal Retrieval Augmented Generation
- SegEarth-R1: Geospatial Pixel Reasoning via Large Language Model
- Evolved Hierarchical Masking for Self-Supervised Learning
- PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
- SDIGLM: Leveraging Large Language Models and Multi-Modal Chain of Thought for Structural Damage Identification
- Spatial Audio Processing with Large Language Model on Wearable Devices
- Steering CLIP's vision transformer with sparse autoencoders
- FindAnything: Open-Vocabulary and Object-Centric Mapping for Robot Exploration in Any Environment
- F3Set: Towards Analyzing Fast, Frequent, and Fine-grained Events from Videos
- VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering
- FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
- Mimic In-Context Learning for Multimodal Tasks
- PACT: Pruning and Clustering-Based Token Reduction for Faster Visual Language Models
- MM-IFEngine: Towards Multimodal Instruction Following
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
- VideoExpert: Augmented LLM for Temporal-Sensitive Video Understanding
- How Can Objects Help Video-Language Understanding?
- SF2T: Self-supervised Fragment Finetuning of Video-LLMs for Fine-Grained Understanding
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
- Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
- OmniCaptioner: One Captioner to Rule Them All
- Endowing Embodied Agents with Spatial Reasoning Capabilities for Vision-and-Language Navigation
- Are We Done with Object-Centric Learning?
- V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
- On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
- Measuring Déjà vu Memorization Efficiently
- SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
- PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
- REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
- OrderChain: Towards General Instruct-Tuning for Stimulating the Ordinal Understanding Ability of MLLM
- SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
- Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
- The 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
- SmolVLM: Redefining small and efficient multimodal models
- InteractVLM: 3D Interaction Reasoning from 2D Foundational Models
- Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
- SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding
- TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
Related