Qwen2.5-VL Technical Report
2025/02/19 by Shuai Bai, Bai, Shuai, Keqin Chen +55 · 1 voice · 2,526 citations
Computer Science · Engineering · #Semiconductor Lasers and Optical Devices #cs.CL #cs.CV
paper · pdf · doi:10.48550/arxiv.2502.13923
Abstract
We introduce Qwen2.5-VL, the latest flagship model of Qwen vision-language series, which demonstrates significant advancements in both foundational capabilities and innovative functionalities. Qwen2.5-VL achieves a major leap forward in understanding and interacting with the world through enhanced visual recognition, precise object localization, robust document parsing, and long-video comprehension. A standout feature of Qwen2.5-VL is its ability to localize objects using bounding boxes or points accurately. It provides robust structured data extraction from invoices, forms, and tables, as well as detailed analysis of charts, diagrams, and layouts. To handle complex inputs, Qwen2.5-VL introduces dynamic resolution processing and absolute time encoding, enabling it to process images of varying sizes and videos of extended durations (up to hours) with second-level event localization. This allows the model to natively perceive spatial scales and temporal dynamics without relying on traditional normalization techniques. By training a native dynamic-resolution Vision Transformer (ViT) from scratch and incorporating Window Attention, we reduce computational overhead while maintaining native resolution. As a result, Qwen2.5-VL excels not only in static image and document understanding but also as an interactive visual agent capable of reasoning, tool usage, and task execution in real-world scenarios such as operating computers and mobile devices. Qwen2.5-VL is available in three sizes, addressing diverse use cases from edge AI to high-performance computing. The flagship Qwen2.5-VL-72B model matches state-of-the-art models like GPT-4o and Claude 3.5 Sonnet, particularly excelling in document and diagram understanding. Additionally, Qwen2.5-VL maintains robust linguistic performance, preserving the core language competencies of the Qwen2.5 LLM.
Cited by
- UniLumos: Fast and Unified Image and Video Relighting with Physics-Plausible Feedback
- UCAgents: Unidirectional Convergence for Visual Evidence Anchored Multi-Agent Medical Decision-Making
- MIRA: Multimodal Iterative Reasoning Agent for Image Editing
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- Unison: A Fully Automatic, Task-Universal, and Low-Cost Framework for Unified Understanding and Generation
- On Conditional Stochastic Interpolation for Generative Nonlinear Sufficient Dimension Reduction
- From Generated Human Videos to Physically Plausible Robot Trajectories
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- LongCat-Image Technical Report
- Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding
- Active Perception Agent for Omnimodal Audio-Video Understanding
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- ProGuard: Towards Proactive Multimodal Safeguard
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- Video Understanding: From Geometry and Semantics to Unified Models
- Bridging Cognitive Gap: Hierarchical Description Learning for Artistic Image Aesthetics Assessment
- MindWatcher: Toward Smarter Multimodal Tool-Integrated Reasoning
- NeXT-IMDL: Build Benchmark for NeXT-Generation Image Manipulation Detection & Localization
- CountGD++: Generalized Prompting for Open-World Counting
- ViLaCD-R1: A Vision-Language Framework for Semantic Change Detection in Remote Sensing
- Text-Routed Sparse Mixture-of-Experts Model with Explanation and Temporal Alignment for Multi-Modal Sentiment Analysis
- MM-UAVBench: How Well Do Multimodal Large Language Models See, Think, and Plan in Low-Altitude UAV Scenarios?
- Diffusion Knows Transparency: Repurposing Video Diffusion for Transparent Object Depth and Normal Estimation
- HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
- MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments
- Modality Inflation: Energy Characterization and Optimization Opportunities for MLLM Inference
- Scaling GUI Agents with Visual State Transitions
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
- EmoCtrl: Controllable Emotional Image Content Generation
- VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- SmartSnap: Proactive Evidence Seeking for Self-Verifying Agents
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
- LiveProteinBench: A Contamination-Free Benchmark for Assessing Models' Specialized Capabilities in Protein Science
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- A Pragmatic VLA Foundation Model
- TimePLE: Rethinking Temporal Representation for Video Temporal Grounding
- MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning
- MulRobBench: A Decision-Level Benchmark for Safe and Security-Policy-Compliant Multimodal UAV Agents
- To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace Expansion
- ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather with Observability Awareness
- PathScale-R1: Cross-scale Reasoning for Pathological Image Analysis
- UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling
- Face Age Verification Vulnerabilities Under Simple Appearance Manipulations
- Breaking the Synthetic-Real Domain Shortcut for Training-Free Generative Replay-based Class Incremental Learning
- Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
- Orient Anything V2: Unifying Orientation and Rotation Understanding
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
- LENS: Adaptive Spatio-Temporal Zooming for Keyframe Sampling in Long-Form Videos
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- iSHIFT: Lightweight Slow-Fast GUI Agent with Adaptive Perception
- A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions
- Layering Virtual Try-On
- ID-V2V: Identity-Preserving Video Restylization
- See Less, See Right: Bi-directional Perceptual Shaping For Multimodal Reasoning
- Coordinated Networking for On-Device Agent-Augmented Real-Time Communication
- Visual Token Compression Enhances Robustness of MLLMs
- StepX-Edge: An On-Device UI Vision-Language Model via Architecture-Training-Deployment Co-Design
- RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation
- Enhancing Pathological VLMs with Cross-scale Reasoning
- Detect Before You Leap: Mirage Detection in Vision-Language Models
- RoboProcessBench: Benchmarking Process-Aware Understanding in Vision-Language Robotic Manipulation
- Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding
- Reason Before You Retrieve: Agentic Planning for Multi-modal RAG
- VlogReward: Learning Multi-Dimensional Evaluation for Vlog Editing
- Group Preference Collapse in Personalized Multimodal Large Language Models
- DocAnnot -- Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Why Does Grounding Hurt Medical VQA? Benchmarking, Diagnosis, and Fine-Tuning of Vision-Language Models
- MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- StAR: Segment Anything Reasoner
- MAI-UI Technical Report: Real-World Centric Foundation GUI Agents
- StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
- Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
- EraseLoRA: MLLM-Driven Foreground Exclusion and Background Subtype Aggregation for Dataset-Free Object Removal
- Scene-VLM: Multimodal Video Scene Segmentation via Vision-Language Models
- Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation
- UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
- From Shallow Humor to Metaphor: Towards Label-Free Harmful Meme Detection via LMM Agent Self-Improvement
- LogicLens: Visual-Logical Co-Reasoning for Text-Centric Forgery Analysis
- AstraNav-World: World Model for Foresight Control and Consistency
- Beyond Memorization: A Multi-Modal Ordinal Regression Benchmark to Expose Popularity Bias in Vision-Language Models
- Streaming Video Instruction Tuning
- AndroidLens: Long-latency Evaluation with Nested Sub-targets for Android GUI Agents
- UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters
- Beyond Pixel Simulation: Pathology Image Generation via Diagnostic Semantic Tokens and Prototype Control
- MMSRARec: Summarization and Retrieval Augumented Sequential Recommendation Based on Multimodal Large Language Model
- Benchmarking and Enhancing VLM for Compressed Image Understanding
- Chain-of-Anomaly Thoughts with Large Vision-Language Models
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- VL4Gaze: Unleashing Vision-Language Models for Gaze Following
- Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
- Learning to Reason in 4D: Dynamic Spatial Understanding for Vision Language Models
- UTDesign: A Unified Framework for Stylized Text Editing and Generation in Graphic Design Images
- Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding
- Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
- LoLA: Long Horizon Latent Action Learning for General Robot Manipulation
- Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation
- SpatialTree: How Spatial Abilities Branch Out in MLLMs
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- SemanticGen: Video Generation in Semantic Space
- Visual-Aware CoT: Achieving High-Fidelity Visual Consistency in Unified Models
- From Indoor to Open World: Revealing the Spatial Reasoning Gap in MLLMs
- WorldWarp: Propagating 3D Geometry with Asynchronous Video Diffusion
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal Reasoning
- QuantiPhy: A Quantitative Benchmark Evaluating Physical Reasoning Abilities of Vision-Language Models
- DeliveryBench: Can Agents Earn Profit in Real World?
- EchoTrail-GUI: Building Actionable Memory for GUI Agents via Critic-Guided Self-Exploration
- TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation
- Generative Human-Object Interaction Detection via Differentiable Cognitive Steering of Multi-modal LLMs
- FC-MIR: A Mobile Screen Awareness Framework for Intent-Aware Recommendation based on Frame-Compressed Multimodal Trajectory Reasoning
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- PENDULUM: A Benchmark for Assessing Sycophancy in Multimodal Large Language Models
- 3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
- ESearch-R1: Learning Cost-Aware MLLM Agents for Interactive Embodied Search via Reinforcement Learning
- SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
- OpenView: Empowering MLLMs with Out-of-view VQA
- Adaptive-VoCo: Complexity-Aware Visual Token Compression for Vision-Language Models
- Region-Constraint In-Context Generation for Instructional Video Editing
- Learning Semantic Atomic Skills for Multi-Task Robotic Manipulation
- Embodied4C: Measuring What Matters for Embodied Vision-Language Navigation
- Enabling Disaggregated Multi-Stage MLLM Inference via GPU-Internal Scheduling and Resource Sharing
- GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Xiaomi MiMo-VL-Miloco Technical Report
- RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- Video Detective: Seek Critical Clues Recurrently to Answer Question from Long Videos
- Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
- MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented Generation
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- The World is Your Canvas: Painting Promptable Events with Reference Images, Trajectories, and Text
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- AdaTooler-V: Adaptive Tool-Use for Images and Videos
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- SceneDiff: A Benchmark and Method for Multiview Object Change Detection
- VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
- Multimodal RewardBench 2: Evaluating Omni Reward Models for Interleaved Text and Image
- GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
- PhysBrain: Human Egocentric Data as a Bridge from Vision Language Models to Physical Intelligence
- CitySeeker: How Do VLMS Explore Embodied Urban Navigation With Implicit Human Needs?
- Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
- VenusBench-GD: A Comprehensive Multi-Platform GUI Benchmark for Diverse Grounding Tasks
- Guiding Perception-Reasoning Closer to Human in Blind Image Quality Assessment
- Collaborative Edge-to-Server Inference for Vision-Language Models
- OS-Oracle: A Comprehensive Framework for Cross-Platform GUI Critic Models
- Scaling Spatial Reasoning in MLLMs through Programmatic Data Synthesis
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- MRG-R1: Reinforcement Learning for Clinically Aligned Medical Report Generation
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- City Navigation in the Wild: Exploring Emergent Navigation from Web-Scale Knowledge in MLLMs
- DSO: Direct Steering Optimization for Bias Mitigation
- DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
- Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Large Video Planner Enables Generalizable Robot Control
- An Efficient and Effective Encoder Model for Vision and Language Tasks in the Remote Sensing Domain
- Step-GUI Technical Report
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- Uni-Parser Technical Report
- Evaluating the Capability of Video Question Generation for Expert Knowledge Elicitation
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- EmoCaliber: Advancing Reliable Visual Emotion Comprehension via Confidence Verbalization and Calibration
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
- TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Focus: A Streaming Concentration Architecture for Efficient Vision-Language Models
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
- Ophiuchus: Incentivizing Tool-augmented "Think with Images" for Joint Medical Segmentation, Understanding and Reasoning
- Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
- OmniDrive-R1: Reinforcement-driven Interleaved Multi-modal Chain-of-Thought for Trustworthy Vision-Language Autonomous Driving
- ChartAgent: A Chart Understanding Framework with Tool Integrated Reasoning
- KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
- DTop-p MoE: Sparsity-Controlled Dynamic Top-p MoE for Foundation Model Pre-training
- HERBench: A Benchmark for Multi-Evidence Integration in Video Question Answering
- A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- Directional Textual Inversion for Personalized Text-to-Image Generation
- Grab-3D: Detecting AI-Generated Videos from 3D Geometric Temporal Consistency
- VLCache: Computing 2% Vision Tokens and Reusing 98% for Vision-Language Inference
- LongVie 2: Multimodal Controllable Ultra-Long Video World Model
- Adapting MLLMs for Nuanced Video Retrieval
- Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
- Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10×
- MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
- LINA: Learning INterventions Adaptively for Physical Alignment and Generalization in Diffusion Models
- ADHint: Adaptive Hints with Difficulty Priors for Reinforcement Learning
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
- Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance
- MMDrive: Interactive Scene Understanding Beyond Vision with Multi-representational Fusion
- Spatial-Aware VLA Pretraining through Visual-Physical Alignment from Human Videos
- GTR-Turbo: Merged Checkpoint is Secretly a Free Teacher for Agentic VLM Training
- STAR: STacked AutoRegressive Scheme for Unified Multimodal Learning
- Motus: A Unified Latent Action World Model
- What Happens Next? Next Scene Prediction with a Unified Video Model
- ProServe: Unified Multi-Priority Request Scheduling for LLM Serving
- SPAR: Session-based Pipeline for Adaptive Retrieval on Legacy File Systems
- From Unlearning to UNBRANDING: A Benchmark for Trademark-Safe Text-to-Image Generation
- Lemon: A Unified and Scalable 3D Multimodal Model for Universal Spatial Understanding
- DrivePI: Spatial-aware 4D MLLM for Unified Autonomous Driving Understanding, Perception, Prediction and Planning
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
- FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
- Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space
- StreamingAssistant: Efficient Visual Token Pruning for Accelerating Online Video Understanding
- More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
- ViInfographicVQA: A Benchmark for Single and Multi-image Visual Question Answering on Vietnamese Infographics
- WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
- UniMark: Artificial Intelligence Generated Content Identification Toolkit
- Using GUI Agent for Electronic Design Automation
- Reconstruction as a Bridge for Event-Based Visual Question Answering
- Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- DentalGPT: Incentivizing Multimodal Complex Reasoning in Dentistry
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction
- SoccerMaster: A Vision Foundation Model for Soccer Understanding
- From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
- SpaceDrive: Infusing Spatial Awareness into VLM-based Autonomous Driving
- Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
- AgriGPT-Omni: A Unified Speech-Vision-Text Framework for Multilingual Agricultural Intelligence
- DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
- Causal Reasoning Favors Encoders: On The Limits of Decoder-Only Models
- Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding
- Boosting RL-Based Visual Reasoning with Selective Adversarial Entropy Intervention
- EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs
- MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
- Grounding Everything in Tokens for Multimodal Large Language Models
- Latent Chain-of-Thought World Modeling for End-to-End Driving
- Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
- Mull-Tokens: Modality-Agnostic Latent Thinking
- AVI-Edit: Audio-sync Video Instance Editing with Granularity-Aware Mask Refiner
- Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- UniUGP: Unifying Understanding, Generation, and Planing For End-to-end Autonomous Driving
- IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
- Rethinking Chain-of-Thought Reasoning for Videos
- UrbanNav: Learning Language-Guided Urban Navigation from Web-Scale Human Trajectories
- H2R-Grounder: A Paired-Data-Free Paradigm for Translating Human Interaction Videos into Physically Grounded Robot Videos
- GAIR: GUI Automation via Information-Joint Reasoning and Group Reflection
- Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning
- ARM-Thinker: Reinforcing Multimodal Generative Reward Models with Agentic Tool Use and Visual Reasoning
- AgentComp: From Agentic Reasoning to Compositional Mastery in Text-to-Image Models
- Same Content, Different Answers: Cross-Modal Inconsistency in MLLMs
- Unified Diffusion Transformer for High-fidelity Text-Aware Image Restoration
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale
- A Multi-Robot Platform for Robotic Triage Combining Onboard Sensing and Foundation Models
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Beyond Real Weights: Hypercomplex Representations for Stable Quantization
- Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models
- Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- From Segments to Scenes: Temporal Understanding for Agentic Autonomous Driving via Vision-Language Models
- Age-Inclusive 3D Human Mesh Recovery for Action-Preserving Data Anonymization
- OpenSubject: Leveraging Video-Derived Identity and Diversity Priors for Subject-driven Image Generation and Manipulation
- PAVAS: Physics-Aware Video-to-Audio Synthesis
- VisKnow: Constructing Visual Knowledge Base for Object Understanding
- Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-and-Language Navigation
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- SpatialDreamer: Incentivizing Spatial Reasoning via Active Mental Imagery
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Relational Visual Similarity
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
- Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
- All You Need Are Random Visual Tokens? Demystifying Token Pruning in VLLMs
- ReLaX: Reasoning with Latent Exploration for Large Reasoning Models
- Geo3DVQA: Evaluating Vision-Language Models for 3D Geospatial Reasoning from Aerial Imagery
- Unified Camera Positional Encoding for Controlled Video Generation
- MMRPT: MultiModal Reinforcement Pre-Training via Masked Vision-Dependent Reasoning
- START: Spatial and Textual Learning for Chart Understanding
- Online Segment Any 3D Thing as Instance Tracking
- Think-Reflect-Revise: A Policy-Guided Reflective Framework for Safety Alignment in Large Vision Language Models
- CLARITY: Medical World Model for Guiding Treatment Decisions by Modeling Context-Aware Disease Trajectories in Latent Space
- See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations
- Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation
- JoPano: Unified Panorama Generation via Joint Modeling
- PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
- Decouple to Generalize: Context-First Self-Evolving Learning for Data-Scarce Vision-Language Reasoning
- MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
- Task-Model Alignment: A Simple Path to Generalizable AI-Generated Image Detection
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- The Role of Entropy in Visual Grounding: Analysis and Optimization
- Scaling Zero-Shot Reference-to-Video Generation
- MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding
- VG-Refiner: Towards Tool-Refined Referring Grounded Reasoning via Agentic Reinforcement Learning
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- ReCAD: Reinforcement Learning Enhanced Parametric CAD Model Generation with Vision-Language Models
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
- Distilling Expert Surgical Knowledge: How to train local surgical VLMs for anatomy explanation in Complete Mesocolic Excision
- InverseCrafter: Efficient Video ReCapture as a Latent Domain Inverse Problem
- ClinTutor-R1: Advancing Scalable and Robust One-to-Many Alignment in Clinical Socratic Education
- Training Multi-Image Vision Agents via End2End Reinforcement Learning
- ProPhy: Progressive Physical Alignment for Dynamic World Simulation
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
- The Dynamic Prior: Understanding 3D Structures for Casual Dynamic Videos
- ShaRP: SHAllow-LayeR Pruning for Video Large Language Models Acceleration
- TV2TV: A Unified Framework for Interleaved Language and Video Generation
- SA-IQA: Redefining Image Quality Assessment for Spatial Aesthetics with Multi-Dimensional Rewards
- Visual Reasoning Tracer: Object-Level Grounded Reasoning Benchmark
- 4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
- Geometry-Aware Single-Image 4D Synthesis via Dense Trajectory Generation
- Can Multi-Modal LLMs Provide Live Step-by-Step Task Guidance?
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- Qwen3.5-Omni Technical Report
- LaFiTe: A Generative Latent Field for 3D Native Texturing
- E3AD: An Emotion-Aware Vision-Language-Action Model for Human-Centric End-to-End Autonomous Driving
- Measuring the Unspoken: A Disentanglement Model and Benchmark for Psychological Analysis in the Wild
- Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
- Towards Cross-View Point Correspondence in Vision-Language Models
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- I2I-Bench: A Comprehensive Benchmark Suite for Image-to-Image Editing Models
- Malicious Image Analysis via Vision-Language Segmentation Fusion: Detection, Element, and Location in One-shot
- When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering
- COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence
- VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
- Refaçade: Editing Object with Given Reference Texture
- PhyVLLM: Physics-Guided Video Language Model with Motion-Appearance Disentanglement
- Not All Birds Look The Same: Identity-Preserving Generation For Birds
- dVLM-AD: Enhance Diffusion Vision-Language-Model for Driving via Controllable Reasoning
- StreamEQA: Towards Streaming Video Understanding for Embodied Scenarios
- SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
- ResponsibleRobotBench: Benchmarking Responsible Robot Manipulation using Multi-modal Large Language Models
- MoReGen: Multi-Agent Motion-Reasoning Engine for Code-based Text-to-Video Synthesis
- Jina-VLM: Small Multilingual Vision Language Model
- TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning
- UniComp: Rethinking Video Compression Through Informational Uniqueness
- UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework
- OmniDexVLG: Learning Dexterous Grasp Generation from Vision Language Model-Guided Grasp Semantics, Taxonomy and Functional Affordance
- AdaptVision: Efficient Vision-Language Models via Adaptive Visual Acquisition
- Colon-X: Advancing Intelligent Colonoscopy toward Clinical Reasoning
- ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos
- Rethinking Prompt Design for Inference-time Scaling in Text-to-Visual Generation
- EEA: Exploration-Exploitation Agent for Long Video Understanding
- Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation
- BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents
- ViDiC: Video Difference Captioning
- Seeing through Imagination: Learning Scene Geometry via Implicit Spatial World Modeling
- MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- Contextual Image Attack: How Visual Context Exposes Multimodal Safety Vulnerabilities
- GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm
- Diagnose, Correct, and Learn from Manipulation Failures via Visual Symbols
- SeeNav-Agent: Enhancing Vision-Language Navigation with Visual Prompt and Step-Level Policy Optimization
- RULER-Bench: Probing Rule-based Reasoning Abilities of Next-level Video Generation Models for Vision Foundation Intelligence
- GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- Distill, Forget, Repeat: A Framework for Continual Unlearning in Text-to-Image Diffusion Models
- dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
- YingVideo-MV: Music-Driven Multi-Stage Video Generation
- Vision to Geometry: 3D Spatial Memory for Sequential Embodied MLLM Reasoning and Exploration
- Does Hearing Help Seeing? Investigating Audio-Video Joint Denoising for Video Generation
- GUI Exploration Lab: Enhancing Screen Navigation in Agents via Multi-Turn Reinforcement Learning
- VACoT: Rethinking Visual Data Augmentation with VLMs
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
- ReVSeg: Incentivizing the Reasoning Chain for Video Segmentation with Reinforcement Learning
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- UnicEdit-10M: A Dataset and Benchmark Breaking the Scale-Quality Barrier via Unified Verification for Reasoning-Enriched Edits
- PhyDetEx: Detecting and Explaining the Physical Plausibility of T2V Models
- OpenREAD: Reinforced Open-Ended Reasoning for End-to-End Autonomous Driving with LLM-as-Critic
- CauSight: Learning to Supersense for Visual Causal Discovery
- Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
- GR-RL: Going Dexterous and Precise for Long-Horizon Robotic Manipulation
- StreamGaze: Gaze-Guided Temporal Reasoning and Proactive Understanding in Streaming Videos
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Generative Editing in the Joint Vision-Language Space for Zero-Shot Composed Image Retrieval
- NavForesee: A Unified Vision-Language World Model for Hierarchical Planning and Dual-Horizon Navigation Prediction
- ViRectify: A Challenging Benchmark for Video Reasoning Correction with Multimodal Large Language Models
- FishDetector-R1: Unified MLLM-Based Framework with Reinforcement Fine-Tuning for Weakly Supervised Fish Detection, Segmentation, and Counting
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards
- IC-World: In-Context Generation for Shared World Modeling
- IGen: Scalable Data Generation for Robot Learning from Open-World Images
- ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images
- PhotoFramer: Multi-modal Image Composition Instruction
- Efficient and Scalable Monocular Human-Object Interaction Motion Reconstruction
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- Accelerating Streaming Video Large Language Models via Hierarchical Token Compression
- Med-CMR: A Fine-Grained Benchmark Integrating Visual Evidence and Clinical Logic for Medical Complex Multimodal Reasoning
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
- OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
- Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
- When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
- Debate with Images: Detecting Deceptive Behaviors in Multimodal Large Language Models
- ChartPoint: Guiding MLLMs with Grounding Reflection for Chart Reasoning
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Video-CoM: Interactive Video Reasoning via Chain of Manipulations
- Visual Generation Tuning
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- OctoMed: Data Recipes for State-of-the-Art Multimodal Medical Reasoning
- LatBot: Distilling Universal Latent Actions for Vision-Language-Action Models
- Obstruction reasoning for robotic grasping
- REVEAL: Reasoning-Enhanced Forensic Evidence Analysis for Explainable AI-Generated Image Detection
- ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- MindPower: Enabling Theory-of-Mind Reasoning in VLM-based Embodied Agents
- McSc: Motion-Corrective Preference Alignment for Video Generation with Self-Critic Hierarchical Reasoning
- Visual Puns from Idioms: An Iterative LLM-T2IM-MLLM Framework
- AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
- Captain Safari: A World Engine with Pose-Aligned 3D Memory
- Ovis-Image Technical Report
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator Optimization
- Seeing before Observable: Potential Risk Reasoning in Autonomous Driving via Vision Language Models
- AgroCoT: A Chain-of-Thought Benchmark for Evaluating Reasoning in Vision-Language Models for Agriculture
- Test-time scaling of diffusions with flow maps
- AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model
- ReasonEdit: Towards Reasoning-Enhanced Image Editing Models
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior
- RoadSceneBench: A Lightweight Benchmark for Mid-Level Road Scene Understanding
- BINDER: Instantly Adaptive Mobile Manipulation with Open-Vocabulary Commands
- Asking like Socrates: Socrates helps VLMs understand remote sensing images
- Agentic Learner with Grow-and-Refine Multimodal Semantic Memory
- Unexplored flaws in multiple-choice VQA evaluations
- UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
- From Compound Figures to Composite Understanding: Developing a Multi-Modal LLM from Biomedical Literature with Medical Multiple-Image Benchmarking and Validation
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- MoGAN: Improving Motion Quality in Video Diffusion via Few-Step Motion Adversarial Post-Training
- VacuumVLA: Boosting VLA Capabilities via a Unified Suction and Gripping Tool for Complex Robotic Manipulation
- REVISOR: Beyond Textual Reflection, Towards Multimodal Introspective Reasoning in Long-Form Video Understanding
- FITRep: Attention-Guided Item Representation via MLLMs
- Thinking With Bounding Boxes: Enhancing Spatio-Temporal Video Grounding via Reinforcement Fine-Tuning
- Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
- Co-Training Vision Language Models for Remote Sensing Multi-task Learning
- SocialNav: Training Human-Inspired Foundation Model for Socially-Aware Embodied Navigation
- CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
- DASIP: Dynamic Test-Time Compute Scaling for Robot Control with Stochastic Interpolant Policies
- Text-Guided Semantic Image Encoder
- Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
- SPHINX: A Synthetic Environment for Visual Perception and Reasoning
- NVIDIA Nemotron Parse 1.1
- A Reason-then-Describe Instruction Interpreter for Controllable Video Generation
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Vision-Language Memory for Spatial Reasoning
- iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation
- Reinforcing Action Policies by Prophesying
- Wanderland: Geometrically Grounded Simulation for Open-World Embodied AI
- The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
- Flash-DMD: Towards High-Fidelity Few-Step Image Generation with Efficient Distillation and Joint Reinforcement Learning
- HBridge: H-Shape Bridging of Heterogeneous Experts for Unified Multimodal Understanding and Generation
- HalDec-Bench: Benchmarking Hallucination Detector in Image Captioning
- MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology
- Look Where It Matters: Training-Free Ultra-HR Remote Sensing VQA via Adaptive Zoom Search
- Object-Centric Vision Token Pruning for Vision Language Models
- Action Without Interaction: Probing the Physical Foundations of Video LMMs via Contact-Release Detection
- Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Towards Benign Memory Forgetting for Selective Multimodal Large Language Model Unlearning
- Vision-Language Models for Automated 3D PET/CT Report Generation
- VICoT-Agent: A Vision-Interleaved Chain-of-Thought Framework for Interpretable Multimodal Reasoning and Scalable Remote Sensing Analysis
- WaymoQA: A Multi-View Visual Question Answering Dataset for Safety-Critical Reasoning in Autonomous Driving
- Arcadia: Toward a Full-Lifecycle Framework for Embodied Lifelong Learning
- Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation
- EmoFeedback2: Reinforcement of Continuous Emotional Image Generation via LVLM-based Reward and Textual Feedback
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- M3Prune: Hierarchical Communication Graph Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
- Reasoning-VLA: A Fast and General Vision-Language-Action Reasoning Model for Autonomous Driving
- MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images
- DiffSeg30k: A Multi-Turn Diffusion Editing Benchmark for Localized AIGC Detection
- CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
- Distilling Counterfactual Reasoning from Language to Vision: Causal Graph Guided Post-Training for Video Understanding
- Syn-GRPO: Self-Evolving Data Synthesis for MLLM Perception Reasoning
- LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
- Advances and Challenges in Solar Flare Prediction: A Review
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
- Scaling Agentic Reinforcement Learning for Tool-Integrated Reasoning in VLMs
- CLASH: A Benchmark for Cross-Modal Contradiction Detection
- CodeV: Code with Images for Faithful Visual Reasoning via Tool-Aware Policy Optimization
- VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
- Be My Eyes: Extending Large Language Models to New Modalities Through Multi-Agent Collaboration
- HunyuanOCR Technical Report
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- VideoChat-M1: Collaborative Policy Planning for Video Understanding via Multi-Agent Reinforcement Learning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- OrdMoE: Preference Alignment via Hierarchical Expert Group Ranking in Multimodal Mixture-of-Experts LLMs
- Table Comprehension in Building Codes using Vision Language Models and Domain-Specific Fine-Tuning
- BackdoorVLM: A Benchmark for Backdoor Attacks on Vision-Language Models
- Parallel Vision Token Scheduling for Fast and Accurate Multimodal LMMs Inference
- VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
- DuoTeach: Dual Role Self-Teaching for Coarse-to-Fine Decision Coordination in Vision--Language Models
- VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-Resolution
- Thinking Ahead: Foresight Intelligence in MLLMs and World Models
- Beyond Description: Cognitively Benchmarking Fine-Grained Action for Embodied Agents
- Towards Efficient VLMs: Information-Theoretic Driven Compression via Adaptive Structural Pruning
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- Disc3D: Automatic Curation of High-Quality 3D Dialog Data via Discriminative Object Referring
- VADE: Variance-Aware Dynamic Sampling via Online Sample-Level Difficulty Estimation for Multimodal RL
- RoadBench: Benchmarking MLLMs on Fine-Grained Spatial Understanding and Reasoning under Urban Road Scenarios
- Introducing Visual Scenes and Reasoning: A More Realistic Benchmark for Spoken Language Understanding
- Mixture of Horizons in Action Chunking
- AutoFocus-IL: VLM-based Saliency Maps for Data-Efficient Visual Imitation Learning without Extra Human Annotations
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal Reasoning
- DocPTBench: Benchmarking End-to-End Photographed Document Parsing and Translation
- ChineseVideoBench: Benchmarking Multi-modal Large Models for Chinese Video Question Answering
- MASS: Motion-Aware Spatial-Temporal Grounding for Physics Reasoning and Comprehension in Vision-Language Models
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- AnyExperts: On-Demand Expert Allocation for Multimodal Language Models with Mixture of Expert
- DiVE-k: Differential Visual Reasoning for Fine-grained Image Recognition
- RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System
- Beyond Words and Pixels: A Benchmark for Implicit World Knowledge Reasoning in Generative Models
- MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
- EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
- ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- EventBench: Towards Comprehensive Benchmarking of Event-based MLLMs
- Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- ARIAL: An Agentic Framework for Document VQA with Precise Answer Localization
- VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
- RAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-Scale
- Pistachio: Towards Synthetic, Balanced, and Long-Form Video Anomaly Benchmarks
- Multi-speaker Attention Alignment for Multimodal Social Interaction
- MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use
- Test-Time Temporal Sampling for Efficient MLLM Video Understanding
- VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- Plan-X: Instruct Video Generation via Semantic Planning
- RynnVLA-002: A Unified Vision-Language-Action and World Model
- Understanding Counting Mechanisms in Large Language and Vision-Language Models
- Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Counterfactual World Models via Digital Twin-conditioned Video Diffusion
- SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
- Beyond Multiple Choice: Verifiable OpenQA for Robust Vision-Language RFT
- Intervene-All-Paths: Unified Mitigation of LVLM Hallucinations across Alignment Formats
- Can MLLMs Read the Room? A Multimodal Benchmark for Assessing Deception in Multi-Party Social Interactions
- VLA-4D: Embedding 4D Awareness into Vision-Language-Action Models for SpatioTemporally Coherent Robotic Manipulation
- FireScope: Wildfire Risk Raster Prediction with a Chain-of-Thought Oracle
- One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
- Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation
- OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- MatPedia: A Universal Generative Foundation for High-Fidelity Material Synthesis
- Q-REAL: Towards Realism and Plausibility Evaluation for AI-Generated Content
- R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
- Pillar-0: A New Frontier for Radiology Foundation Models
- SceneDesigner: Controllable Multi-Object Image Generation with 9-DoF Pose Manipulation
- VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
- Personalized Reward Modeling for Text-to-Image Generation
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Learning to Think Fast and Slow for Visual Language Models
- You Only Forward Once: An Efficient Compositional Judging Paradigm
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- POMA-3D: The Point Map Way to 3D Scene Understanding
- Decoupling Complexity from Scale in Latent Diffusion Model
- "To Survive, I Must Defect": Jailbreaking LLMs via the Game-Theory Scenarios
- Mantis: A Versatile Vision-Language-Action Model with Disentangled Visual Foresight
- UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
- Video2Layout: Recall and Reconstruct Metric-Grounded Cognitive Map for Spatial Reasoning
- Text2Traffic: A Text-to-Image Generation and Editing Method for Traffic Scenes
- Generative Photographic Control for Scene-Consistent Video Cinematic Editing
- DeepSport: A Multimodal Large Language Model for Comprehensive Sports Video Reasoning via Agentic Reinforcement Learning
- When to Think and When to Look: Uncertainty-Guided Lookback
- CrossCheck-Bench: Diagnosing Compositional Failures in Multimodal Conflict Resolution
- ChartEditor: A Reinforcement Learning Framework for Robust Chart Editing
- Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
- SplitFlux: Learning to Decouple Content and Style from a Single Image
- UniSER: A Foundation Model for Unified Soft Effects Removal
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- Reasoning via Video: The First Evaluation of Video Models' Reasoning Abilities through Maze-Solving Tasks
- MoDES: Accelerating Mixture-of-Experts Multimodal Large Language Models via Dynamic Expert Skipping
- Attention Grounded Enhancement for Visual Document Retrieval
- OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
- Let Language Constrain Geometry: Vision-Language Models as Semantic and Spatial Critics for 3D Generation
- ManipShield: A Unified Framework for Image Manipulation Detection, Localization and Explanation
- AsyncVLA: Asynchronous Flow Matching for Vision-Language-Action Models
- Jailbreaking Large Vision Language Models in Intelligent Transportation Systems
- GRPO Privacy Is at Risk: A Membership Inference Attack Against Reinforcement Learning With Verifiable Rewards
- RoboTidy : A 3D Gaussian Splatting Household Tidying Benchmark for Embodied Navigation and Action
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image Restoration
- Training-Free Multi-View Extension of IC-Light for Textual Position-Aware Scene Relighting
- PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
- P1: Mastering Physics Olympiads with Reinforcement Learning
- SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design
- Video Finetuning Improves Reasoning Between Frames
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- PIGEON: VLM-Driven Object Navigation via Points of Interest Selection
- Sensing and Understanding the World over Air: A Large Multimodal Model for Mobile Networks
- ViSS-R1: Self-Supervised Reinforcement Video Reasoning
- From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Direct Visual Grounding by Directing Attention of Visual Tokens
- EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
- MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- Reasoning Text-to-Video Retrieval via Digital Twin Video Representations and Large Language Models
- Fast Reasoning Segmentation for Images and Videos
- Constructing and Interpreting Digital Twin Representations for Visual Reasoning via Reinforcement Learning
- Explainable AI-Generated Image Detection RewardBench
- CrossVid: A Comprehensive Benchmark for Evaluating Cross-Video Reasoning in Multimodal Large Language Models
- PRISM of Opinions: A Persona-Reasoned Multimodal Framework for User-centric Conversational Stance Detection
- Learning to Hear by Seeing: It's Time for Vision Language Models to Understand Artistic Emotion from Sight and Sound
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- KVSwap: Disk-aware KV Cache Offloading for Long-Context On-device Inference
- TopoPerception: A Shortcut-Free Evaluation of Global Visual Perception in Large Vision-Language Models
- Sat2RealCity: Geometry-Aware and Appearance-Controllable 3D Urban Generation from Satellite Imagery
- Hindsight Distillation Reasoning with Knowledge Encouragement Preference for Knowledge-based Visual Question Answering
- VIDEOP2R: Video Understanding from Perception to Reasoning
- Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- VisMem: Latent Vision Memory Unlocks Potential of Vision-Language Models
- Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis for Large Reasoning Models
- EgoEMS: A High-Fidelity Multimodal Egocentric Dataset for Cognitive Assistance in Emergency Medical Services
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- Enhancing the Outcome Reward-based RL Training of MLLMs with Self-Consistency Sampling
- MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation
- Towards Trustworthy Dermatology MLLMs: A Benchmark and Multimodal Evaluator for Diagnostic Narratives
- MACEval: A Multi-Agent Continual Evaluation Network for Large Models
- TIP and Polish: Text-Image-Prototype Guided Multi-Modal Generation via Commonality-Discrepancy Modeling and Refinement
- OSGym: Scalable OS Infra for Computer Use Agents
- | \circlearrowright \boxedBUS |: A Large and Diverse Multimodal Benchmark for evaluating the ability of Vision-Language Models to understand Rebus Puzzles
- MARC: Multimodal and Multi-Task Agentic Retrieval-Augmented Generation for Cold-Start Recommender System
- NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
- Why does weak-OOD help? A Further Step Towards Understanding Jailbreaking VLMs
- Where and What Matters: Sensitivity-Aware Task Vectors for Many-Shot Multimodal In-Context Learning
- An Efficient Training Pipeline for Reasoning Graphical User Interface Agents
- Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
- Neurophysiological Characteristics of Adaptive Reasoning for Creative Problem-Solving Strategy
- From Exploration to Exploitation: A Two-Stage Entropy RLVR Approach for Noise-Tolerant MLLM Training
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Grounding Computer Use Agents on Human Demonstrations
- Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
- PanoNav: Mapless Zero-Shot Object Navigation with Panoramic Scene Parsing and Dynamic Memory
- SRNN: Spatiotemporal Relational Neural Network for Intuitive Physics Understanding
- How Do VLAs Effectively Inherit from VLMs?
- MVU-Eval: Towards Multi-Video Understanding Evaluation for Multimodal LLMs
- JPRO: Automated Multimodal Jailbreaking via Multi-Agent Collaboration Framework
- NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
- TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
- SportR: A Benchmark for Multimodal Large Language Model Reasoning in Sports
- TimeSense:Making Large Language Models Proficient in Time-Series Analysis
- ReMoD: Rethinking Modality Contribution in Multimodal Stance Detection via Dual Reasoning
- SARCH: Multimodal Search for Archaeological Archives
- Culture in Action: Evaluating Text-to-Image Models through Social Activities
- LiveStar: Live Streaming Assistant for Real-World Online Video Understanding
- Pressure2Motion: Hierarchical Human Motion Reconstruction from Ground Pressure with Text Guidance
- Visual Spatial Tuning
- Tracking and Understanding Object Transformations
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- V-Thinker: Interactive Thinking with Images
- T-FIX: Text-Based Explanations with Features Interpretable to eXperts
- GUI-360^∘: A Comprehensive Dataset and Benchmark for Computer-Using Agents
- Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment
- Context informs pragmatic interpretation in vision-language models
- What's in Common? Multimodal Models Hallucinate When Reasoning Across Scenes
- SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
- SurgAnt-ViVQA: Learning to Anticipate Surgical Events through GRU-Driven Temporal Cross-Attention
- QG-CoC: Question-Guided Chain-of-Captions for Large Multimodal Models
- ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- Can Visual Input Be Compressed? A Visual Token Compression Benchmark for Large Multimodal Models
- Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation
- SAIL-RL: Guiding MLLMs in When and How to Think via Dual-Reward RL Tuning
- When Modalities Conflict: How Unimodal Reasoning Uncertainty Governs Preference Dynamics in MLLMs
- When Visualizing is the First Step to Reasoning: MIRA, a Benchmark for Visual Chain-of-Thought
- Towards Selection of Large Multimodal Models as Engines for Burned-in Protected Health Information Detection in Medical Images
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- A Unified Reasoning Framework for Holistic Zero-Shot Video Anomaly Analysis
- RefTon: Reference person shot assist virtual Try-on
- GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
- CueBench: Advancing Unified Understanding of Context-Aware Video Anomalies in Real-World
- ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
- Saliency-R1: Incentivizing Unified Saliency Reasoning Capability in MLLM with Confidence-Guided Reinforcement Learning
- VinciCoder: Unifying Multimodal Code Generation via Coarse-to-fine Visual Reinforcement Learning
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Text-guided Fine-Grained Video Anomaly Understanding
- iFlyBot-VLA Technical Report
- PreferThinker: Reasoning-based Personalized Image Preference Assessment
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
- Multi-Modal Feature Fusion for Spatial Morphology Analysis of Traditional Villages via Hierarchical Graph Neural Networks
- HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
- GUI-Rise: Structured Reasoning and History Summarization for GUI Navigation
- Cognitive Alignment in Personality Reasoning: Leveraging Prototype Theory for MBTI Inference
- LongCat-Flash-Omni Technical Report
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning
- Toward Accurate Long-Horizon Robotic Manipulation: Language-to-Action with Foundation Models via Scene Graphs
- Towards Universal Video Retrieval: Generalizing Video Embedding via Synthesized Multimodal Pyramid Curriculum
- NAUTILUS: A Large Multimodal Model for Underwater Scene Understanding
- A Multi-Modal Neuro-Symbolic Approach for Spatial Reasoning-Based Visual Grounding in Robotics
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- ChartAB: A Benchmark for Chart Grounding & Dense Alignment
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- Emu3.5: Native Multimodal Models are World Learners
- RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration
- Envisioning Future Interactive Web Development: Editing Webpage with Natural Language
- Counteracting Matthew Effect in Self-Improvement of LVLMs through Head-Tail Re-balancing
- LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation
- CRAG-MM: Multi-modal Multi-turn Comprehensive RAG Benchmark
- EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Dynamic VLM-Guided Negative Prompting for Diffusion Models
- ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
- MedVLSynther: Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- PairUni: Pairwise Training for Unified Multimodal Language Models
- ALDEN: Reinforcement Learning for Active Navigation and Evidence Gathering in Long Documents
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
- FlowMM: Cross-Modal Information Flow Guided KV Cache Merging for Efficient Multimodal Context Inference
- MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding
- SpatialQ: Understanding 3D Gaussian Splatting Scene Quality via Visual-based MLLM
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models
- TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
- Work Zones challenge VLM Trajectory Planning: Toward Mitigation and Robust Autonomous Driving
- Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
- Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction
- Progressive Multimodal Alignment for Continual Instruction Tuning
- Unlimited OCR Works
- CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering
- Seeking the Unfamiliar but Memorable: Conceptual Creativity as Meta-Learning
- FlowInOne:Unifying Multimodal Generation as Image-in, Image-out Flow Matching
- Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- RL makes MLLMs see better than SFT
- MentisOculi: Revealing the Limits of Reasoning with Mental Imagery
- Better Call Grep: Evaluating and Improving Grep-Like Lexical Retrieval for Repository-Level Code Completion
- Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models
- Input Domain Aware MoE: Decoupling Routing Decisions from Task Optimization in Mixture of Experts
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Metis-SPECS: Decoupling Multimodal Learning via Self-distilled Preference-based Cold Start
- SmoothGuard: Defending Multimodal Large Language Models with Noise Perturbation and Clustering Aggregation
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing Guidance
- Group Relative Attention Guidance for Image Editing
- OSWorld-MCP: Benchmarking MCP Tool Invocation In Computer-Use Agents
- Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
- Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
- REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Model
- MGA: Memory-Driven GUI Agent for Observation-Centric Interaction
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- Enhancing Vision-Language Models for Autonomous Driving through Task-Specific Prompting and Spatial Reasoning
- Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- Pie: A Programmable Serving System for Emerging LLM Applications
- Advancing Off-Road Autonomous Driving: The Large-Scale ORAD-3D Dataset and Comprehensive Benchmarks
- TeleEgo: Benchmarking Egocentric AI Assistants in the Wild
- Conflict Adaptation in Vision-Language Models
- Reasoning Visual Language Model for Chest X-Ray Analysis
- World Simulation with Video Foundation Models for Physical AI
- Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
- Victim as a Service: Designing a System for Engaging with Interactive Scammers
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- Track, Inpaint, Resplat: Subject-driven 3D and 4D Generation with Progressive Texture Infilling
- PixelRefer: A Unified Framework for Spatio-Temporal Object Referring with Arbitrary Granularity
- PRISM-Bench: A Benchmark of Puzzle-Based Visual Tasks with CoT Error Detection
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT
- M4FC: a Multimodal, Multilingual, Multicultural, Multitask Real-World Fact-Checking Dataset
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
- Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
- VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
- Revisiting Multimodal Positional Encoding in Vision-Language Models
- Evaluation of Vision-LLMs in Surveillance Video
- Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- Multi-Stage Field Extraction of Financial Documents with OCR and Compact Vision-Language Models
- LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
- Positional Preservation Embedding for Multimodal Large Language Models
- Understanding What Is Not Said:Referring Remote Sensing Image Segmentation with Scarce Expressions
- REVISION:Reflective Intent Mining and Online Reasoning Auxiliary for E-commerce Visual Search System Optimization
- IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- RoboSVG: A Unified Framework for Interactive SVG Generation with Multi-modal Guidance
- OFFSIDE: Benchmarking Unlearning Misinformation in Multimodal Large Language Models
- Jarvis: Towards Personalized AI Assistant via Personal KV-Cache Retrieval
- Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents
- DynaSolidGeo: A Dynamic Benchmark for Genuine Spatial Mathematical Reasoning of VLMs in Solid Geometry
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- LongCat-Video Technical Report
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation
- LightAgent: Mobile Agentic Foundation Models
- Head Pursuit: Probing Attention Specialization in Multimodal Transformers
- Bridging the gap to real-world language-grounded visual concept learning
- TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
- Towards Fine-Grained Human Motion Video Captioning
- NoisyGRPO: Incentivizing Multimodal CoT Reasoning via Noise Injection and Bayesian Estimation
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
- 3DReasonKnee: Advancing Grounded Reasoning in Medical Vision Language Models
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- GeoThought: A Dataset for Enhancing Mathematical Geometry Reasoning in Vision-Language Models
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- GhostEI-Bench: Do Mobile Agents Resilience to Environmental Injection in Dynamic On-Device Environments?
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- Discovering Intersectional Bias via Directional Alignment in Face Recognition Embeddings
- SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
- Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
- LayerComposer: Multi-Human Personalized Generation via Layered Canvas
- Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning
- Metis-HOME: Hybrid Optimized Mixture-of-Experts for Multimodal Reasoning
- Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets
- EQPO: Equitable Group Relative Policy Optimization for Clinical Reasoning
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- Vision-language models learn the geometry of human perceptual space
- From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction
- MedReason-R1: Learning to Reason for CT Diagnosis with Reinforcement Learning and Local Zoom
- Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
- MINED: Probing and Updating with Multimodal Time-Sensitive Knowledge for Large Multimodal Models
- Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
- AutoMT: A Multi-Agent LLM Framework for Automated Metamorphic Testing of Autonomous Driving Systems
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- ColorAgent: Building A Robust, Personalized, and Interactive OS Agent
- DaMo: Data Mixing Optimizer in Fine-tuning Multimodal LLMs for Mobile Phone Agents
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Structured and Abstractive Reasoning on Multi-modal Relational Knowledge Images
- Preliminary Use of Vision Language Model Driven Extraction of Mouse Behavior Towards Understanding Fear Expression
- Exploring Scale Shift in Crowd Localization under the Context of Domain Generalization
- olmOCR 2: Unit Test Rewards for Document OCR
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- See the Text: From Tokenization to Visual Reading
- VDRive: Leveraging Reinforced VLA and Diffusion Policy for End-to-end Autonomous Driving
- OCR-Quality: A Human-Annotated Dataset for OCR Quality Assessment
- CUARewardBench: A Benchmark for Evaluating Reward Models on Computer-using Agent
- Embodied Navigation with Auxiliary Task of Action Description Prediction
- UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
- Proactive Reasoning-with-Retrieval Framework for Medical Multimodal Large Language Models
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
- VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- DSI-Bench: A Benchmark for Dynamic Spatial Intelligence
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- MoGA: Mixture-of-Groups Attention for End-to-End Long Video Generation
- UniGenBench++: A Unified Semantic Evaluation Benchmark for Text-to-Image Generation
- DeepSeek-OCR: Contexts Optical Compression
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- HouseTour: A Virtual Real Estate A(I)gent
- Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- VERA-V: Variational Inference Framework for Jailbreaking Vision-Language Models
- MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues
- UniRL-Zero: Reinforcement Learning on Unified Models with Joint Language Model and Diffusion Model Experts
- PICABench: How Far Are We from Physically Realistic Image Editing?
- Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
- iDETEX: Empowering MLLMs for Intelligent DETailed EXplainable IQA
- SimpleVSF: VLM-Scoring Fusion for Trajectory Prediction of End-to-End Autonomous Driving
- See or Say Graphs: Agent-Driven Scalable Graph Structure Understanding with Vision-Language Models
- ZSPAPrune: Zero-Shot Prompt-Aware Token Pruning for Vision-Language Models
- Robobench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- Symmetric Entropy-Constrained Video Coding for Machines
- UltraCUA: A Foundation Model for Computer Use Agents with Hybrid Action
- Video Reasoning without Training
- Training-free Online Video Step Grounding
- Fly-CL: A Fly-Inspired Framework for Enhancing Efficient Decorrelation and Reduced Training Time in Pre-trained Model-based Continual Representation Learning
- Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
- Segmentation as A Plug-and-Play Capability for Frozen Multimodal LLMs
- Pursuing Minimal Sufficiency in Spatial Reasoning
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- BLIP3o-NEXT: Next Frontier of Native Image Generation
- ImagerySearch: Adaptive Test-Time Search for Video Generation Beyond Semantic Dependency Constraints
- Select Less, Reason More: Prioritizing Evidence Purity for Video Reasoning
- GUIrilla: A Scalable Framework for Automated Desktop UI Exploration
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Constantly Improving Image Models Need Constantly Improving Benchmarks
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
- Free-Grained Hierarchical Recognition
- xLLM Technical Report
- In-Context Learning with Unpaired Clips for Instruction-based Video Editing
- ColorBench: Benchmarking Mobile Agents with Graph-Structured Framework for Complex Long-Horizon Tasks
- Hi-Agent: Hierarchical Vision-Language Agents for Mobile Device Control
- AI for Service: Proactive Assistance with AI Glasses
- Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
- Identity-GRPO: Optimizing Multi-Human Identity-preserving Video Generation via Reinforcement Learning
- Kelle: Co-design KV Caching and eDRAM for Efficient LLM Serving in Edge Computing
- Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
- TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
- Multimodal Function Vectors for Visual Relations
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- RoboHiMan: A Hierarchical Evaluation Paradigm for Compositional Generalization in Long-Horizon Manipulation
- DriveCritic: Towards Context-Aware, Human-Aligned Evaluation for Autonomous Driving with Vision-Language Models
- Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- Generative Universal Verifier as Multimodal Meta-Reasoner
- Reasoning in Space via Grounding in the World
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- Detect Anything via Next Point Prediction
- Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
- K-frames: Scene-Driven Any-k Keyframe Selection for long video understanding
- VideoLucy: Deep Memory Backtracking for Long Video Understanding
- An Empirical Study for Representations of Videos in Video Question Answering via MLLMs
- SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
- HackWorld: Evaluating Computer-Use Agents on Exploiting Web Application Vulnerabilities
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
- SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models
- HoneyBee: Data Recipes for Vision-Language Reasoners
- DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving
- Scaling Language-Centric Omnimodal Representation Learning
- IVEBench: Modern Benchmark Suite for Instruction-Guided Video Editing Assessment
- ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- VideoNSA: Native Sparse Attention Scales Video Understanding
- VidGuard-R1: AI-Generated Video Detection and Explanation via Reasoning MLLMs and RL
- ODI-Bench: Can MLLMs Understand Immersive Omnidirectional Environments?
- LikePhys: Evaluating Intuitive Physics Understanding in Video Diffusion Models via Likelihood Preference
- DocReward: A Document Reward Model for Structuring and Stylizing
- Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
- InternSVG: Towards Unified SVG Tasks with Multimodal Large Language Models
- Vision-LLMs for Spatiotemporal Traffic Forecasting
- CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- Demystifying Numerosity in Diffusion Models -- Limitations and Remedies
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- GeoVLMath: Enhancing Geometry Reasoning in Vision-Language Models via Cross-Modal Reward for Auxiliary Line Creation
- A Survey on Agentic Multimodal Large Language Models
- Video-STR: Reinforcing MLLMs in Video Spatio-Temporal Reasoning with Relation Graph
- Judge Before Answer: Can MLLM Discern the False Premise in Question?
- More than A Point: Capturing Uncertainty with Adaptive Affordance Heatmaps for Spatial Grounding in Robotic Tasks
- CodePlot-CoT: Mathematical Visual Reasoning by Thinking with Code-Driven Images
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- MatterDoor: Sampling Zero-shot Spatio-semantic Priors using Generative Models
- R-WoM: Retrieval-augmented World Model For Computer-use Agents
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- Unlocking Vision-Language Models for Video Anomaly Detection via Fine-Grained Prompting
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- OmniQuality-R: Advancing Reward Models Through All-Encompassing Quality Assessment
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
- VR-Thinker: Boosting Video Reward Models through Thinking-with-Image Reasoning
- OpusAnimation: Code-Based Dynamic Chart Generation
- Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- TCMA: Text-Conditioned Multi-granularity Alignment for Drone Cross-Modal Text-Video Retrieval
- CompassNav: Steering From Path Imitation To Decision Understanding In Navigation
- Think Twice to See More: Iterative Visual Reasoning in Medical VLMs
- Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
- Answer-Consistent Chain-of-thought Reinforcement Learning For Multi-modal Large Langauge Models
- From Generic to Specialized: A Subspecialty Diagnostic System Powered by Self-Supervised Learning for Cervical Histopathology
- Reallocating Attention Across Layers to Reduce Multimodal Hallucination
- Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning
- Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
- Known By Their Actions: Fingerprinting LLM Browser Agents via UI Traces
- Asymmetric Flow Models
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- AutoPR: Let's Automate Your Academic Promotion!
- Spotlight on Token Perception for Multimodal Reinforcement Learning
- MomentSeg: Moment-Centric Sampling for Enhanced Video Pixel Understanding
- Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning
- Towards Efficient Multimodal Unified Reasoning Model via Model Merging
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- Multimodal Policy Internalization for Conversational Agents
- PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
- Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
- Vision Language Models: A Survey of 26K Papers
- Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
- Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- InstructX: Towards Unified Visual Editing with MLLM Guidance
- ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping
- UniVideo: Unified Understanding, Generation, and Editing for Videos
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- Kelp: A Streaming Safeguard for Large Models via Latent Dynamics-Guided Risk Detection
- A Multimodal Depth-Aware Method For Embodied Reference Understanding
- NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
- A Large-scale Dataset for Robust Complex Anime Scene Text Detection
- MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- Towards Proprioception-Aware Embodied Planning for Dual-Arm Humanoid Robots
- GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- IntentionVLA: Generalizable and Efficient Embodied Intention Reasoning for Human-Robot Interaction
- Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- MoA-VR: A Mixture-of-Agents System Towards All-in-One Video Restoration
- CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
- RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
- Beyond Textual CoT: Interleaved Text-Image Chains with Deep Confidence Reasoning for Image Editing
- SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- TALENT: Table VQA via Augmented Language-Enhanced Natural-text Transcription
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- SaFeR-VLM: Toward Safety-aware Fine-grained Reasoning in Multimodal Models
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- StaR-KVQA: Structured Reasoning Traces for Implicit-Knowledge Visual Question Answering
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- What MLLMs Learn about When they Learn about Multimodal Reasoning
- VLA-R1: Enhancing Reasoning in Vision-Language-Action Models
- Get RICH or Die Scaling: Profitably Trading Inference Compute for Robustness
- TalkCuts: A Large-Scale Dataset for Multi-Shot Human Speech Video Generation
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding
- HOI-R1: Exploring the Potential of Multimodal Large Language Models for Human-Object Interaction Detection
- FORGE-Tree: Diffusion-Forcing Tree Search for Long-Horizon Robot Manipulation
- Presenting a Paper is an Art: Self-Improvement Aesthetic Agents for Academic Presentations
- MADIAVE: Multi-Agent Debate for Implicit Attribute Value Extraction
- When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- Video-LMM Post-Training: A Deep Dive into Video Reasoning with Large Multimodal Models
- Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
- Say One Thing, Do Another? Diagnosing Reasoning-Execution Gaps in VLM-Powered Mobile-Use Agents
- Pack and Force Your Memory: Long-form and Consistent Video Generation
- ContextNav: Towards Agentic Multimodal In-Context Learning
- TBStar-Edit: From Image Editing Pattern Shifting to Consistency Enhancement
- VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- Contrastive Representation Regularization for Vision-Language-Action Models
- Debunk and Infer: Multimodal Fake News Detection via Diffusion-Generated Evidence and LLM Reasoning
- Human Behavior Atlas: Benchmarking Unified Psychological and Social Behavior Understanding
- Watch and Learn: Learning to Use Computers from Online Videos
- Pulp Motion: Framing-aware multimodal camera and human motion generation
- A.I.R.: Enabling Adaptive, Iterative, and Reasoning-based Frame Selection For Video Question Answering
- ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
- AlphaApollo: A System for Deep Agentic Reasoning
- Turning Drift into Constraint: Robust Reasoning Alignment in Non-Stationary Multi-Stream Environments
- Fine-Grained GRPO for Precise Preference Alignment in Flow Models
- RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- AgriGPT-VL: Agricultural Vision-Language Understanding Suite
- RetiBridge: Bridging Quantitative Retinal Biomarkers and Qualitative Diagnosis with a Knowledge-Guided Multimodal Large Language Model
- GUI-Spotlight: Adaptive Iterative Focus Refinement for Enhanced GUI Visual Grounding
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- MLLMEraser: Achieving Test-Time Unlearning in Multimodal Large Language Models through Activation Steering
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine Differentiation
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- FrameOracle: Learning What to See and How Much to See in Videos
- UniShield: An Adaptive Multi-Agent Framework for Unified Forgery Image Detection and Localization
- MonitorVLM:A Vision Language Framework for Safety Violation Detection in Mining Operations
- Cross-Modal Content Optimization for Steering Web Agent Preferences
- NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
- GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- SpineBench: A Clinically Salient, Level-Aware Benchmark Powered by the SpineMed-450k Corpus
- Towards Scalable and Consistent 3D Editing
- TIT-Score: Evaluating Long-Prompt Based Text-to-Image Alignment via Text-to-Image-to-Text Consistency
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- MoGIC: Boosting Motion Generation via Intention Understanding and Visual Context
- One Patch to Caption Them All: A Unified Zero-Shot Captioning Framework
- Cache-to-Cache: Direct Semantic Communication Between Large Language Models
- Image Generation Based on Image Style Extraction
- IMAGEdit: Let Any Subject Transform
- Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- EditTrack: Detecting and Attributing AI-assisted Image Editing
- GEM: A Gym for Agentic LLMs
- ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- It Takes Two: Your GRPO Is Secretly DPO
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- ACPO: Adaptive Curriculum Policy Optimization for Aligning Vision-Language Models in Complex Reasoning
- LVLMs as inspectors: an agentic framework for category-level structural defect annotation
- Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
- Plug-and-Play Prompt Refinement via Latent Feedback for Diffusion Model Alignment
- PAL-UI: Planning with Active Look-back for Vision-Based GUI Agents
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- RiskPO: Risk-based Policy Optimization via Verifiable Reward for LLM Post-Training
- BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
- MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
- Disentangling Foreground and Background for vision-Language Navigation via Online Augmentation
- Apriel-1.5-15b-Thinker
- TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding
- Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
- AccidentBench: Benchmarking Multimodal Understanding and Reasoning in Vehicle Accidents and Beyond
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Stable Cinemetrics : Structured Taxonomy and Evaluation for Professional Video Generation
- Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
- SCUBA: Salesforce Computer Use Benchmark
- PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
- MR2-Bench: Going Beyond Matching to Reasoning in Multimodal Retrieval
- IMG: Calibrating Diffusion Models via Implicit Multimodal Guidance
- Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
- EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
- Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
- DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
- Reinforced Embodied Planning with Verifiable Reward for Real-World Robotic Manipulation
- MuSLR: Multimodal Symbolic Logical Reasoning
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
- Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- LaTo: Landmark-tokenized Diffusion Transformer for Fine-grained Human Face Editing
- Importance Sampling for Multi-Negative Multimodal Direct Preference Optimization
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- DescribeEarth: Describe Anything for Remote Sensing Images
- OceanGym: A Benchmark Environment for Underwater Embodied Agents
- Automated Model Discovery via Multi-modal & Multi-step Pipeline
- Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- LLaVAShield: Safeguarding Multimodal Multi-Turn Dialogues in Vision-Language Models
- Probing the Limits of Stylistic Alignment in Vision-Language Models
- Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play
- Seeing Before Reasoning: A Unified Framework for Generalizable and Explainable Fake Image Detection
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- PixelCraft: A Multi-Agent System for High-Fidelity Visual Reasoning on Structured Images
- VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
- OIG-Bench: A Multi-Agent Annotated Benchmark for Multimodal One-Image Guides Understanding
- RealUnify: Do Unified Models Truly Benefit from Unification? A Comprehensive Benchmark
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- ZOO-Prune: Training-Free Token Pruning via Zeroth-Order Gradient Estimation in Vision-Language Models
- LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
- IWR-Bench: Can LVLMs reconstruct interactive webpage from a user interaction video?
- AdaDetectGPT: Adaptive Detection of LLM-Generated Text with Statistical Guarantees
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- OnomatoGen: Onomatopoeia Generation with the Alpha-Channel in Manga
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Latent Visual Reasoning
- UniVid: The Open-Source Unified Video Model
- NeMo: Needle in a Montage for Video-Language Understanding
- Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO
- DepthLM: Metric Depth From Vision Language Models
- DocPruner: A Storage-Efficient Framework for Multi-Vector Visual Document Retrieval via Adaptive Patch-Level Embedding Pruning
- Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
- Euclid's Gift: Enhancing Spatial Perception and Reasoning in Vision-Language Models via Geometric Surrogate Tasks
- FreeRet: MLLMs as Training-Free Retrievers
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- GeoVLM-R1: Reinforcement Fine-Tuning for Improved Remote Sensing Reasoning
- Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
- VisualOverload: Probing Visual Understanding of VLMs in Really Dense Scenes
- ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
- Assessing Large Language Models in Updating Their Forecasts with New Information
- HiViS: Hiding Visual Tokens from the Drafter for Speculative Decoding in Vision-Language Models
- EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
- Towards Fine-Grained Text-to-3D Quality Assessment: A Benchmark and A Two-Stage Rank-Learning Metric
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- Falcon: A Cross-Modal Evaluation Dataset for Comprehensive Safety Perception
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- Poivre: Self-Refining Visual Pointing with Reinforcement Learning
- GUI-Shepherd: Reliable Process Reward and Verification for Long-Sequence GUI Tasks
- HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling
- LUQ: Layerwise Ultra-Low Bit Quantization for Multimodal Large Language Models
- Video Panels for Long Video Understanding
- HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection
- ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- Efficient Multi-turn RL for GUI Agents via Decoupled Training and Adaptive Data Curation
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
- Dynamic-TreeRPO: Breaking the Independent Trajectory Bottleneck with Structured Sampling
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- Decoupling Reasoning and Perception: An LLM-LMM Framework for Faithful Visual Reasoning
- GUI-PRA: Process Reward Agent for GUI Tasks
- Self-Consistency as a Free Lunch: Reducing Hallucinations in Vision-Language Models via Self-Reflection
- Culture In a Frame: C3B as a Comic-Based Benchmark for Multimodal Culturally Awareness
- BEV-VLM: Trajectory Planning via Unified BEV Abstraction
- LAGEA: Language Guided Embodied Agents for Robotic Manipulation
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Planning with Unified Multimodal Models
- Open-Vocabulary Spatio-Temporal Scene Graph for Robot Perception and Teleoperation Planning
- MMPB: It's Time for Multi-Modal Personalization
- VideoScore2: Think before You Score in Generative Video Evaluation
- CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
- InterActHuman: Multi-Concept Human Animation with Layout-Aligned Audio Conditions
- Learn the Ropes, Then Trust the Wins: Self-imitation with Progressive Exploration for Agentic Reinforcement Learning
- JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- REMA: A Unified Reasoning Manifold Framework for Interpreting Large Language Model
- Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation
- ReLAM: Learning Anticipation Model for Rewarding Visual Robotic Manipulation
- RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
- From Watch to Imagine: Steering Long-horizon Manipulation via Human Demonstration and Future Envisionment
- MimicDreamer: Aligning Human and Robot Demonstrations for Scalable VLA Training
- REFINE-CONTROL: A Semi-supervised Distillation Method For Conditional Image Generation
- SciTS: Scientific Time Series Understanding and Generation with LLMs
- Lightweight Structured Multimodal Reasoning for Clinical Scene Understanding in Robotics
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- From Bias to Balance: Exploring and Mitigating Spatial Bias in LVLMs
- RISK: A Framework for GUI Agents in E-commerce Risk Management
- Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
- MultiCrafter: High-Fidelity Multi-Subject Generation via Disentangled Attention and Identity-Aware Preference Alignment
- Perception-Consistency Multimodal Large Language Models Reasoning via Caption-Regularized Policy Optimization
- D-Artemis: A Deliberative Cognitive Framework for Mobile GUI Multi-Agents
- Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Training-Free Multimodal Deepfake Detection via Graph Reasoning
- CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
- Guiding Audio Editing with Audio Language Model
- CompareBench: A Benchmark for Visual Comparison Reasoning in Vision-Language Models
- Learning GUI Grounding with Spatial Reasoning from Visual Feedback
- VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
- VideoChat-R1.5: Visual Test-Time Scaling to Reinforce Multimodal Reasoning by Iterative Perception
- GeoRef: Referring Expressions in Geometry via Task Formulation, Synthetic Supervision, and Reinforced MLLM-based Solutions
- Autoregressive End-to-End Planning with Time-Invariant Spatial Alignment and Multi-Objective Policy Refinement
- SPADE: Structured Pruning and Adaptive Distillation for Efficient LLM-TTS
- DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning
- MTRDrive: Memory-Tool Synergistic Reasoning for Robust Autonomous Driving in Corner Cases
- Meta-Memory: Retrieving and Integrating Semantic-Spatial Memories for Robot Spatial Reasoning
- Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- SafeSteer: Adaptive Subspace Steering for Efficient Jailbreak Defense in Vision-Language Models
- RAD: Towards Trustworthy Retrieval-Augmented Multi-modal Clinical Diagnosis
- OmniScene: Attention-Augmented Multimodal 4D Scene Understanding for Autonomous Driving
- Logics-Parsing Technical Report
- FreezeVLA: Action-Freezing Attacks against Vision-Language-Action Models
- MAD: Motion Appearance Decoupling for efficient Driving World Models
- MMHBench: A Multi-Perspective Benchmark for Mental Health Understanding in Long-Form Videos
- JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles
- Objective-Aligned Direct Answer SFT for Robust Multi-Frame Medical VQA
- LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
- FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification
- IndustryForge-27B: A Domain-Enhanced Multimodal Foundation Model for Industrial CAD
- VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- Flux-OPD: On-Policy Distillation with Evolving Contexts
- Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments
- SleepLM: Natural-Language Intelligence for Human Sleep
- See What You Need: Query-Aware Visual Intelligence through Reasoning-Perception Loops
- UniAPO: Unified Multimodal Automated Prompt Optimization
- Instant Preference Alignment for Text-to-Image Diffusion Models
- Proximal Supervised Fine-Tuning
- ConViS-Bench: Estimating Video Similarity Through Semantic Concepts
- Citrus-V: Advancing Medical Foundation Models with Unified Medical Image Grounding for Clinical Reasoning
- Pure Vision Language Action (VLA) Models: A Comprehensive Survey
- How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
- MAPO: Mixed Advantage Policy Optimization
- MVT: Mask-Grounded Vision-Language Models for Taxonomy-Aligned Land-Cover Tagging
- Understanding-in-Generation: Reinforcing Generative Capability of Unified Model via Infusing Understanding into Generation
- OraPO: Oracle-educated Reinforcement Learning for Data-efficient and Factual Radiology Report Generation
- The 1st Solution for MOSEv2 Challenge 2025: Long-term and Concept-aware Video Segmentation via SeC
- OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps
- Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Bi-VLM: Pushing Ultra-Low Precision Post-Training Quantization Boundaries in Vision-Language Models
- UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
- Multi-Agent Amodal Completion: Direct Synthesis with Fine-Grained Semantic Guidance
- SD-VLM: Spatial Measuring and Understanding with Depth-Encoded Vision-Language Models
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Table2LaTeX-RL: High-Fidelity LaTeX Code Generation from Table Images via Reinforced Multimodal Language Models
- Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
- LoT-Pass: Long-term-robust Image Watermarking for Image to Video Generation
- RoboSeek: You Need to Interact with Your Objects
- NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
- Qwen3-Omni Technical Report
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
- AgriDoctor: A Multimodal Intelligent Assistant for Agriculture
- The 1st Solution for 7th LSVOS RVOS Track: SaSaSa2VA
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
- Eye Gaze Tells You Where to Compute: Gaze-Driven Efficient VLMs
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
- CoReVLA: A Dual-Stage End-to-End Autonomous Driving Framework for Long-Tail Scenarios via Collect-and-Refine
- Vision-Language Models as Differentiable Semantic and Spatial Rewards for Text-to-3D Generation
- Qianfan-VL: Domain-Enhanced Universal Vision-Language Models
- SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models
- TennisTV: Do Multimodal Large Language Models Understand Tennis Rallies?
- ChartMaster: Advancing Chart-to-Code Generation with Real-World Charts and Chart Similarity Reinforcement Learning
- BTL-UI: Blink-Think-Link Reasoning Model for GUI Agent
- GUI-ARP: Enhancing Grounding with Adaptive Region Perception for GUI Agents
- Lynx: Towards High-Fidelity Personalized Video Generation
- BaseReward: A Strong Baseline for Multimodal Reward Model
- Structured Information for Improving Spatial Relationships in Text-to-Image Generation
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- The Iconicity of the Generated Image
- RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation
- Generalizable Geometric Image Caption Synthesis
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
- How Good are Foundation Models in Step-by-Step Embodied Reasoning?
- MedFact-R1: Towards Factual Medical Reasoning via Pseudo-Label Augmentation
- AutoEdit: Automatic Hyperparameter Tuning for Image Editing
- T2I-ReasonBench: Benchmarking Reasoning-Informed Text-to-Image Generation
- Embodied Arena: A Comprehensive, Unified, and Evolving Evaluation Platform for Embodied AI
- Frame Sampling Strategies Matter: A Benchmark for small vision language models
- MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks
- TriSPrompt: A Hierarchical Soft Prompt Model for Multimodal Rumor Detection with Incomplete Modalities
- A Multi-To-One Interview Paradigm for Efficient MLLM Evaluation
- TableDART: Dynamic Adaptive Multi-Modal Routing for Table Understanding
- GenExam: A Multidisciplinary Text-to-Image Exam
- TGPO: Tree-Guided Preference Optimization for Robust Web Agent Reinforcement Learning
- Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
- Wan-Animate: Unified Character Animation and Replacement with Holistic Replication
- Can Current AI Models Count What We Mean, Not What They See? A Benchmark and Systematic Evaluation
- ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding
- Structures Meet Semantics: Multimodal Fusion via Graph Contrastive Learning
- SAIL-VL2 Technical Report
- Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
- LLM-I: LLMs are Naturally Interleaved Multimodal Creators
- See, Think, Act: Teaching Multimodal Agents to Effectively Interact with GUI by Identifying Toggles
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Mind the (Language) Gap: Towards Probing Numerical and Cross-Lingual Limits of LVLMs
- MEENA (PersianMMMU): Multimodal-Multilingual Educational Exams for N-level Assessment
- EdiVal-Agent: An Object-Centric Framework for Automated, Fine-Grained Evaluation of Multi-Turn Editing
- The LLM Already Knows: Estimating LLM-Perceived Question Difficulty via Hidden Representations
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Lego-Edit: A General Image Editing Framework with Model-Level Bricks and MLLM Builder
- 3D Aware Region Prompted Vision Language Model
- ActiveVLN: Towards Active Exploration via Multi-Turn RL in Vision-and-Language Navigation
- EfficientUICoder: Efficient MLLM-based UI Code Generation via Input and Output Token Compression
- Embodied Navigation Foundation Model
- EgoMem: Lifelong Memory Agent for Full-duplex Omnimodal Models
- Dr.V: A Hierarchical Perception-Temporal-Cognition Framework to Diagnose Video Hallucination by Fine-grained Spatial-Temporal Grounding
- NeuroStrike: Neuron-Level Attacks on Aligned LLMs
- Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
- Igniting VLMs toward the Embodied Space
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- UI-S1: Advancing GUI Automation via Semi-online Reinforcement Learning
- ConEQsA: Concurrent and Asynchronous Embodied Questions Scheduling and Answering
- When Safe Unimodal Inputs Collide: Optimizing Reasoning Chains for Cross-Modal Safety in Multimodal Large Language Models
- Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
- Adapting and Evaluating Multimodal Large Language Models for Adolescent Idiopathic Scoliosis Self-Management: A Divide and Conquer Framework
- VideoAgent: Personalized Synthesis of Scientific Videos
- DreamNav: A Trajectory-Based Imaginative Framework for Zero-Shot Vision-and-Language Navigation
- Action Hints: Semantic Typicality and Context Uniqueness for Generalizable Skeleton-based Video Anomaly Detection
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Environmental Injection Attacks against GUI Agents in Realistic Dynamic Environments
- ReFineG: Synergizing Small Supervised Models and LLMs for Low-Resource Grounded Multimodal NER
- Detecting Text Manipulation in Images using Vision Language Models
- MagicMirror: A Large-Scale Dataset and Benchmark for Fine-Grained Artifacts Assessment in Text-to-Image Generation
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Towards Understanding Visual Grounding in Visual Language Models
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Unified Multimodal Model as Auto-Encoder
- Measuring Epistemic Humility in Multimodal Large Language Models
- Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation Synthesis
- Visual Programmability: A Guide for Code-as-Thought in Chart Understanding
- DATE: Dynamic Absolute Time Enhancement for Long Video Understanding
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- VQualA 2025 Challenge on Visual Quality Comparison for Large Multimodal Models: Methods and Results
- OpenFake: An Open Dataset and Platform Toward Real-World Deepfake Detection
- How well can LLMs provide planning feedback in grounded environments?
- AdsQA: Towards Advertisement Video Understanding
- MobileRL: Online Agentic Reinforcement Learning for Mobile GUI Agents
- Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- Visual Representation Alignment for Multimodal Large Language Models
- Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- TA-VLA: Elucidating the Design Space of Torque-aware Vision-Language-Action Models
- TextlessRAG: End-to-End Visual Document RAG by Speech Without Text
- In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting
- GLEAM: Learning to Match and Explain in Cross-View Geo-Localization
- Video Parallel Scaling: Aggregating Diverse Frame Subsets for VideoLLMs
- KLIPA: A Knowledge Graph and LLM-Driven QA Framework for IP Analysis
- Reconstruction Alignment Improves Unified Multimodal Models
- Aligning Large Vision-Language Models by Deep Reinforcement Learning and Direct Preference Optimization
- Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- Interleaving Reasoning for Better Text-to-Image Generation
- Light-Weight Cross-Modal Enhancement Method with Benchmark Construction for UAV-based Open-Vocabulary Object Detection
- Multimodal Reasoning for Science: Technical Report and 1st Place Solution to the ICML 2025 SeePhys Challenge
- PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters
- SparkUI-Parser: Enhancing GUI Perception with Robust Grounding and Parsing
- COMMET: A System for Human-Induced Conflicts in Mobile Manipulation of Everyday Tasks
- Learning Active Perception via Self-Evolving Preference Optimization for GUI Grounding
- FPC-VLA: A Vision-Language-Action Framework with a Supervisor for Failure Prediction and Correction
- ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
- ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
- Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- FLM-Audio: Natural Monologues Improves Native Full-Duplex Chatbots via Dual Training
- TeRA: Rethinking Text-guided Realistic 3D Avatar Generation
- OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- CMRAG: Co-modality-based visual document retrieval and question answering
- Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- AutoDrive-R2: Incentivizing Reasoning and Self-Reflection Capacity for VLA Model in Autonomous Driving
- PixFoundation 2.0: Do Video Multi-Modal LLMs Use Motion in Visual Grounding?
- RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
- Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering
- Kwai Keye-VL 1.5 Technical Report
- Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Less Redundancy: Boosting Practicality of Vision Language Model in Walking Assistants
- InterPose: Learning to Generate Human-Object Interactions from Large-Scale Web Videos
- Prompt the Unseen: Evaluating Visual-Language Alignment Beyond Supervision
- LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- Galaxea Open-World Dataset and G0 Dual-System VLA Model
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- The Demon is in Ambiguity: Revisiting Situation Recognition with Single Positive Multi-Label Learning
- Is this chart lying to me? Automating the detection of misleading visualizations
- ELV-Halluc: Benchmarking Semantic Aggregation Hallucinations in Long Video Understanding
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- UItron: Foundational GUI Agent with Advanced Perception and Planning
- R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
- Estimating 2D Keypoints of Surgical Tools Using Vision-Language Models with Low-Rank Adaptation
- Intern-S1: A Scientific Multimodal Foundation Model
- Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
- Waver: Wave Your Way to Lifelike Video Generation
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- OneReward: Unified Mask-Guided Image Generation via Multi-Task Human Preference Learning
- Veritas: Generalizable Deepfake Detection via Pattern-Aware Reasoning
- AudioStory: Generating Long-Form Narrative Audio with Large Language Models
- 11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
- SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
- PG-Agent: An Agent Powered by Page Graph
- InquireMobile: Teaching VLM-based Mobile Agent to Request Human Assistance via Reinforcement Fine-Tuning
- Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
- CVBench: Evaluating Cross-Video Synergies for Complex Multimodal Understanding and Reasoning
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- Enhancing Document VQA Models via Retrieval-Augmented Generation
- CrossHOI-Bench: A Unified Benchmark for HOI Evaluation across Vision-Language Models and HOI-Specific Methods
- The Mind's Eye: A Multi-Faceted Reward Framework for Guiding Visual Metaphor Generation
- Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models
- Wan-S2V: Audio-Driven Cinematic Video Generation
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
- Object Detection with Multimodal Large Vision-Language Models: An In-depth Review
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- MedRepBench: A Comprehensive Benchmark for Medical Report Interpretation
- SentiMM: A Multimodal Multi-Agent Framework for Sentiment Analysis in Social Media
- SurgWound-Bench: A Benchmark for Surgical Wound Diagnosis
- Mobile-Agent-v3: Fundamental Agents for GUI Automation
- Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
- CyPortQA: Benchmarking Multimodal Large Language Models for Cyclone Preparedness in Port Operation
- MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation
- MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models
- HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
- OmniTry: Virtual Try-On Anything without Masks
- AdaDocVQA: Adaptive Framework for Long Document Visual Question Answering in Low-Resource Settings
- Breaking the SFT Plateau: Multimodal Structured Reinforcement Learning for Chart-to-Code Generation
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- Structured Prompting and Multi-Agent Knowledge Distillation for Traffic Video Interpretation and Risk Inference
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- LENS: Learning to Segment Anything with Unified Reinforced Reasoning
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- Precise Action-to-Video Generation Through Visual Action Prompts
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- "DIVE" into Hydrogen Storage Materials Discovery with AI Agents
- HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
- Creative4U: MLLMs-based Advertising Creative Image Selector with Comparative Reasoning
- DianJin-OCR-R1: Enhancing OCR Capabilities via a Reasoning-and-Tool Interleaved Vision-Language Model
- Lumen: Consistent Video Relighting and Harmonious Background Replacement with Video Generative Models
- Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
- Adaptive Reinforcement for Open-ended Medical Reasoning via Semantic-Guided Reward Collapse Mitigation
- MIRAGE: Towards AI-Generated Image Detection in the Wild
- RadarQA: Multi-modal Quality Analysis of Weather Radar Forecasts
- You Don't Know Until You Click:Automated GUI Testing for Production-Ready Software Evaluation
- Temporal Grounding as a Learning Signal for Referring Video Object Segmentation
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- OmniD: Generalizable Robot Manipulation Policy via Image-Based BEV Representation
- Thyme: Think Beyond Images
- Language models align with brain regions that represent concepts across modalities
- Better Supervised Fine-tuning for VQA: Integer-Only Loss
- Ovis2.5 Technical Report
- CRAFT-GUI: Curriculum-Reinforced Agent For GUI Tasks
- Labels or Input? Rethinking Augmentation in Multimodal Hate Detection
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
- AEGIS: Authenticity Evaluation Benchmark for AI-Generated Video Sequences
- EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
- NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
- ChatENV: An Interactive Vision-Language Model for Sensor-Guided Environmental Monitoring and Scenario Simulation
- Vision Generalist Model: A Survey
- A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
- We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- Empowering Multimodal LLMs with External Tools: A Comprehensive Survey
- Few-shot Vision-based Human Activity Recognition with MLLM-based Visual Reinforcement Learning
- Contrast Sensitivity in Multimodal Large Language Models: A Psychophysics-Inspired Evaluation
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- GoViG: Goal-Conditioned Visual Navigation Instruction Generation
- Episodic Memory Representation for Long-form Video Understanding
- IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model
- Beyond Blanket Masking: Examining Granularity for Privacy Protection in Images Captured by Blind and Low Vision Users
- OpenCUA: Open Foundations for Computer-Use Agents
- SMA: Who Said That? Auditing Membership Leakage in Semi-Black-box RAG Controlling
- Utilizing Multilingual Encoders to Improve Large Language Models for Low-Resource Languages
- CARES: Collaborative Agentic Reasoning for Error Detection in Surgery
- OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
- SafeFix: Targeted Model Repair via Controlled Image Generation
- Yan: Foundational Interactive Video Generation
- DocThinker: Explainable Multimodal Large Language Models with Rule-based Reinforcement Learning for Document Understanding
- Bridging Formal Language with Chain-of-Thought Reasoning to Geometry Problem Solving
- Safe Semantics, Unsafe Interpretations: Tackling Implicit Reasoning Safety in Large Vision-Language Models
- Towards Effective MLLM Jailbreaking Through Balanced On-Topicness and OOD-Intensity
- ODYSSEY: Open-World Quadrupeds Exploration and Manipulation for Long-Horizon Tasks
- TBAC-UniImage: Unified Understanding and Generation by Ladder-Side Diffusion Tuning
- Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models
- MME-Emotion: A Holistic Evaluation Benchmark for Emotional Intelligence in Multimodal Large Language Models
- FineBadminton: A Multi-Level Dataset for Fine-Grained Badminton Video Understanding
- Matrix-3D: Omnidirectional Explorable 3D World Generation
- LET-US: Long Event-Text Understanding of Scenes
- Understanding Dynamic Scenes in Ego Centric 4D Point Clouds
- Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
- Find Them All: Unveiling MLLMs for Versatile Person Re-identification
- MeteorPred: A Meteorological Multimodal Large Model and Dataset for Severe Weather Event Prediction
- MDK12-Bench: A Comprehensive Evaluation of Multimodal Large Language Models on Multidisciplinary Exams
- Effective Training Data Synthesis for Improving MLLM Chart Understanding
- Mediator-Guided Multi-Agent Collaboration among Open-Source Models for Medical Decision-Making
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models
- Can Large Models Fool the Eye? A New Turing Test for Biological Animation
- Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models
- LoRA in LoRA: Towards Parameter-Efficient Architecture Expansion for Continual Visual Instruction Tuning
- WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent
- Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- MOSEv2: A More Challenging Dataset for Video Object Segmentation in Complex Scenes
- Shuffle-R1: Efficient RL framework for Multimodal Large Language Models via Data-centric Dynamic Shuffle
- Uni-cot: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
- Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
- IAD-R1: Reinforcing Consistent Reasoning in Industrial Anomaly Detection
- Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction
- A Metric for MLLM Alignment in Large-scale Recommendation
- G-UBS: Towards Robust Understanding of Implicit Feedback via Group-Aware User Behavior Simulation
- InfiGUI-G1: Advancing GUI Grounding with Adaptive Exploration Policy Optimization
- Multimodal LLM-assisted Evolutionary Search for Programmatic Control Policies
- SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
- FinMMR: Make Financial Numerical Reasoning More Multimodal, Comprehensive, and Challenging
- NavA3: Understanding Any Instruction, Navigating Anywhere, Finding Anything
- Training-Free Multimodal Large Language Model Orchestration
- ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning
- GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- PET2Rep: Towards Vision-Language Model-Drived Automated Radiology Report Generation for Positron Emission Tomography
- Static and Plugged: Make Embodied Evaluation Simple
- Do Vision-Language Models Leak What They Learn? Adaptive Token-Weighted Model Inversion Attacks
- Beyond the Visible: Benchmarking Occlusion Perception in Multimodal Large Language Models
- VeriGUI: Verifiable Long-Chain GUI Dataset
- Uncertainty-Aware GUI Agent: Adaptive Perception through Component Recommendation and Human-in-the-Loop Refinement
- Are We on the Right Way for Assessing Document Retrieval-Augmented Generation?
- Refining Critical Thinking in LLM Code Generation: A Faulty Premise-based Evaluation Framework
- SAM2-UNeXT: An Improved High-Resolution Baseline for Adapting Foundation Models to Downstream Segmentation Tasks
- Quality-Aware Language-Conditioned Local Auto-Regressive Anomaly Synthesis and Detection
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Towards Trustworthy Multimodal Moderation via Policy-Aligned Reasoning and Hierarchical Labeling
- Pay What LLM Wants: Can LLM Simulate Economics Experiment with 522 Real-human Persona?
- Geoint-R1: Formalizing Multimodal Geometric Reasoning with Dynamic Auxiliary Constructions
- AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
- LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
- MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
- Refine-IQA: Multi-Stage Reinforcement Finetuning for Perceptual Image Quality Assessment
- Following Route Instructions using Large Vision-Language Models: A Comparison between Low-level and Panoramic Action Spaces
- A Multi-Agent System for Complex Reasoning in Radiology Visual Question Answering
- DreamVVT: Mastering Realistic Video Virtual Try-On in the Wild via a Stage-Wise Diffusion Transformer Framework
- Qwen-Image Technical Report
- Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- MedVLThinker: Simple Baselines for Multimodal Medical Reasoning
- Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
- MonoDream: Monocular Vision-Language Navigation with Panoramic Dreaming
- Engagement Prediction of Short Videos with Large Multimodal Models
- Intention-Guided Cognitive Reasoning for Egocentric Long-Term Action Anticipation
- A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models
- Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- What Makes "Good" Distractors for Object Hallucination Evaluation in Large Vision-Language Models?
- Web-CogReasoner: Towards Multimodal Knowledge-Induced Cognitive Reasoning for Web Agents
- MAP: Mitigating Hallucinations in Large Vision-Language Models with Map-Level Attention Processing
- Simulated Ensemble Attack: Transferring Jailbreaks Across Fine-tuned Vision-Language Models
- RoboMemory: A Brain-inspired Multi-memory Agentic Framework for Interactive Environmental Learning in Physical Embodied Systems
- SpectrumWorld: Artificial Intelligence Foundation for Spectroscopy
- PiKV: KV Cache Management System for Mixture of Experts
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- iSafetyBench: A video-language benchmark for safety in industrial environment
- CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
- DocTron-Formula: Generalized Formula Recognition in Complex and Structured Scenarios
- Diagnostic Accuracy of Open-Source Vision-Language Models on Diverse Medical Imaging Tasks
- CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning
- Text-Aware Image Restoration with Diffusion Models
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning
- X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
- MMAT-1M: A Large Reasoning Dataset for Multimodal Agent Tuning
- ArtSeek: Deep artwork understanding via multimodal in-context reasoning and late interaction retrieval
- CAPE: A CLIP-Aware Pointing Ensemble of Complementary Heatmap Cues for Embodied Reference Understanding
- Exploring the Link Between Bayesian Inference and Embodied Intelligence: Toward Open Physical-World Embodied AI Systems
- The Evolution of Video Anomaly Detection: A Unified Framework from DNN to MLLM
- GR-3 Technical Report
- Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning
- Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval
- ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- Multimodal LLMs as Customized Reward Models for Text-to-Image Generation
- Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak Supervision
- ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
- RIS-LAD: A Benchmark and Model for Referring Low-Altitude Drone Image Segmentation
- Geometric-Mean Policy Optimization
- T2I-Copilot: A Training-Free Multi-Agent Text-to-Image System for Enhanced Prompt Interpretation and Interactive Generation
- Enhancing Spatial Reasoning through Visual and Textual Thinking
- GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
- LRR-Bench: Left, Right or Rotate? Vision-Language models Still Struggle With Spatial Understanding Tasks
- The Policy Cliff: A Theoretical Analysis of Reward-Policy Maps in Large Language Models
- A Survey of Token Compression for Efficient Multimodal Large Language Models
- Position: Reasoning After Perception Means Reasoning Without Vision
- CircuitProbe: Tracing Visual Temporal Evidence Flow in Video Language Models
- Object-centric Video Question Answering with Visual Grounding and Referring
- MMBench-GUI: Hierarchical Multi-Platform Evaluation Framework for GUI Agents
- OctoNav: Towards Generalist Embodied Navigation
- OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth?
- RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
- LoViC: Efficient Long Video Generation with Context Compression
- IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
- Datasets and Recipes for Video Temporal Grounding via Reinforcement Learning
- Chart-R1: Chain-of-Thought Supervision and Reinforcement for Advanced Chart Reasoner
- MobileUse: A GUI Agent with Hierarchical Reflection for Autonomous Mobile Operation
- DiagR1: A Vision-Language Model Trained via Reinforcement Learning for Digestive Pathology Diagnosis
- Critique of impure reason: Unveiling the reasoning behaviour of medical large language models
- InsightX Agent: An LMM-based Agentic Framework with Integrated Tools for Reliable X-ray NDT Analysis
- Seeing is Free, Speaking is Not: Uncovering the True Energy Bottleneck in Edge VLM Inference
- U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs
- PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation
- Qwen-Image-2.0 Technical Report
- BusterX++: Towards Unified Cross-Modal AI-Generated Content Detection and Explanation with MLLM
- MagicGUI: A Foundational Mobile GUI Agent with Scalable Data Pipeline and Reinforcement Fine-tuning
- ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
- Benchmarking Gaslighting Negation Attacks Against Reasoning Models
- WebGuard: Building a Generalizable Guardrail for Web Agents
- Constructing Ophthalmic MLLM for Positioning-diagnosis Collaboration Through Clinical Cognitive Chain Reasoning
- A Versatile Pathology Co-pilot via Reasoning Enhanced Multimodal Large Language Model
- Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution
- ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning
- Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
- Pixels to Principles: Probing Intuitive Physics Understanding in Multimodal Language Models
- C2-Evo: Co-Evolving Multimodal Data and Model for Self-Improving Reasoning
- ReasonVQA: A Multi-hop Reasoning Benchmark with Structural Knowledge for Visual Question Answering
- SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting
- AD2-Bench: A Hierarchical CoT Benchmark for MLLM in Autonomous Driving under Adverse Conditions
- A High-Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
- DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
- Describe Anything Model for Visual Question Answering on Text-rich Images
- Mitigating Object Hallucinations via Sentence-Level Early Intervention
- Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
- Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy
- VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video Generation
- Safeguarding Multimodal Knowledge Copyright in the RAG-as-a-Service Environment
- NavComposer: Composing Language Instructions for Navigation Trajectories through Action-Scene-Object Modularization
- UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
- ExpliCIT-QA: Explainable Code-Based Image Table Question Answering
- Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
- Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
- EmbRACE-3K: Embodied Reasoning and Action in Complex Environments
- SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
- Draft-based Approximate Inference for LLMs
- ExpStar: Towards Automatic Commentary Generation for Multi-discipline Scientific Experiments
- VIKI-R: Coordinating Embodied Multi-Agent Cooperation via Reinforcement Learning
- Ambiguity-Aware and High-Order Relation Learning for Multi-Grained Image-Text Matching
- PPJudge: Towards Human-Aligned Assessment of Artistic Painting Process
- AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
- Prompt4Trust: A Reinforcement Learning Prompt Augmentation Framework for Clinically-Aligned Confidence Calibration in Multimodal Large Language Models
- Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
- Multilingual Multimodal Software Developer for Code Generation
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
- Beyond the Linear Separability Ceiling: Aligning Representations in VLMs
- StarDojo: Benchmarking Open-Ended Behaviors of Agentic Multimodal LLMs in Production-Living Simulations with Stardew Valley
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- AI Should Sense Better, Not Just Scale Bigger: Adaptive Sensing as a Paradigm Shift
- PyVision: Agentic Vision with Dynamic Tooling
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- The Synergy Dilemma of Long-CoT SFT and RL: Investigating Post-Training Techniques for Reasoning VLMs
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- Spatial ModernBERT: Spatial-Aware Transformer for Table and Key-Value Extraction in Financial Documents at Scale
- Adjudicated Captioning: Multi-Agent Alignment Scoring and Consensus-Distilled Beam Arbitration for Strict Zero-Shot Image Captioning
- Comprehensive Evaluation of Large Multimodal Models for Nutrition Analysis: A New Benchmark Enriched with Contextual Metadata
- Is Diversity All You Need for Scalable Robotic Manipulation?
- NeoBabel: A Multilingual Open Tower for Visual Generation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- MLlm-DR: Towards Explainable Depression Recognition with MultiModal Large Language Models
- MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment
- GTA1: GUI Test-time Scaling Agent
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Semantic Frame Interpolation
- HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
- Reinforcement Fine-Tuning Naturally Mitigates Forgetting in Continual Post-Training
- Neural-Driven Image Editing
- Is It Time for the Renaissance of Salient Object Detection in the Era of MLLMs?
- Can Prompt Difficulty be Online Predicted for Accelerating RL Finetuning of Reasoning Models?
- CoT-lized Diffusion: Let's Reinforce T2I Generation Step-by-step
- "Hi AirStar, Guide Me to the Badminton Court."
- Multi-Modal Semantic Parsing for the Interpretation of Tombstone Inscriptions
- EmoPrefer: Can Large Language Models Understand Human Emotion Preferences?
- ZERO: Industry-ready Vision Foundation Model with Multi-modal Prompts
- TopoMAS: Large Language Model Driven Topological Materials Multiagent System
- Animation Needs Attention: A Holistic Approach to Slides Animation Comprehension with Visual-Language Models
- A Comparative Study of Specialized LLMs as Dense Retrievers
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- ReSum: Synergizing LLM Reasoning and Summarization with Reinforcement Learning
- BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset
- VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
- Multimodal Mathematical Reasoning with Diverse Solving Perspective
- TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- Kwai Keye-VL Technical Report
- IC-Custom: Diverse Image Customization via In-Context Learning
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- Activation Reward Models for Few-Shot Model Alignment
- ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving
- How Do Vision-Language Models Process Conflicting Information Across Modalities?
- AVC-DPO: Aligned Video Captioning via Direct Preference Optimization
- CI-VID: A Coherent Interleaved Text-Video Dataset
- MedGround-R1: Advancing Medical Image Grounding via Spatial-Semantic Rewarded Group Relative Policy Optimization
- Temporal Chain of Thought: Long-Video Understanding by Thinking in Frames
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- CAVALRY-V: A Large-Scale Generator Framework for Adversarial Attacks on Video MLLMs
- Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
- Box-QAymo: Box-Referring VQA Dataset for Autonomous Driving
- M3MAD-Bench: Multi-Dimensional Evaluation of Multi-Agent Debate Across Domains and Modalities
- Just Noticeable Difference for Large Multimodal Models
- VisioMath: Benchmarking Figure-based Mathematical Reasoning in LMMs
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints
- SCORP: Scene-Consistent Object Refinement via Proxy Generation and Tuning
- PAC Bench: Do Foundation Models Understand Prerequisites for Executing Manipulation Policies?
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video Generation
- Unified Multimodal Understanding via Byte-Pair Visual Encoding
- MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
- Why Reinforcement Fine-Tuning Enables MLLMs Preserve Prior Knowledge Better: A Data Perspective
- A Survey on Vision-Language-Action Models for Autonomous Driving
- ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
- Proteus-ID: ID-Consistent and Motion-Coherent Video Customization
- Holistic Artificial Intelligence in Medicine; improved performance and explainability
- GeoProg3D: Compositional Visual Reasoning for City-Scale 3D Language Fields
- From Individuals to Interactions: Benchmarking Gender Bias in Multimodal Large Language Models from the Lens of Social Relationship
- CyberV: Cybernetics for Test-time Scaling in Video Understanding
- Empowering Small VLMs to Think with Dynamic Memorization and Exploration
- IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
- Ovis-U1 Technical Report
- SoMi-ToM: Evaluating Multi-Perspective Theory of Mind in Embodied Social Interactions
- Iterative Zoom-In: Temporal Interval Exploration for Long Video Understanding
- DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
- MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning
- Seg-R1: Segmentation Can Be Surprisingly Simple with Reinforcement Learning
- HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains
- Looking Beyond Visible Cues: Implicit Video Question Answering via Dual-Clue Reasoning
- Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
- EFRame: Deeper Reasoning via Exploration-Filter-Replay Reinforcement Learning Framework
- Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
- Visual Structures Helps Visual Reasoning: Addressing the Binding Problem in VLMs
- Universal Retrieval for Multimodal Trajectory Modeling
- SPAZER: Spatial-Semantic Progressive Reasoning Agent for Zero-shot 3D Visual Grounding
- R1-Track: Direct Application of MLLMs to Visual Object Tracking via Reinforcement Learning
- MiCo: Multi-image Contrast for Reinforcement Visual Reasoning
- ImplicitQA: Going beyond frames towards Implicit Video Reasoning
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
- Super Encoding Network: Recursive Association of Multi-Modal Encoders for Video Understanding
- ReME: A Data-Centric Framework for Training-Free Open-Vocabulary Segmentation
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- BitVLA: 1-bit Vision-Language-Action Models for Robotics Manipulation
- Team PA-VCG's Solution for Competition on Understanding Chinese College Entrance Exam Papers in ICDAR'25
- Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance
- FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- DeepVideo-R1: Video Reinforcement Fine-Tuning via Difficulty-aware Regressive GRPO
- FOCUS: Internal MLLM Representations for Efficient Fine-Grained Visual Question Answering
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models
- Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
- MMSearch-R1: Incentivizing LMMs to Search
- OR-VSKC: Resolving Visual-Semantic Knowledge Conflicts in Operating Rooms with Synthetic Data-Guided Alignment
- Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
- UniCode2: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
- Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
- Fine-grained Token Allocation Via Operation Pruning for Efficient MLLMs
- Fake or Real, Can Robots Tell? Evaluating VLM Robustness to Domain Shift in Single-View Robotic Scene Understanding
- BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models
- Doc2SAR: A Synergistic Framework for High-Fidelity Extraction of Structure-Activity Relationships from Scientific Documents
- Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification
- GTR-CoT: Graph Traversal as Visual Chain of Thought for Molecular Structure Recognition
- Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
- Surgery-R1: Advancing Surgical-VQLA with Reasoning Multimodal Large Language Model via Reinforcement Learning
- ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- Holmes: Towards Effective and Harmless Model Ownership Verification to Personalized Large Vision Models via Decoupling Common Features
- jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval
- Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
- Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
- OmniGen2: Towards Instruction-Aligned Multimodal Generation
- Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset
- Event-Priori-Based Vision-Language Model for Efficient Visual Understanding
- GUI-Reflection: Empowering Multimodal GUI Models with Self-Reflection Behavior
- SUDER: Self-Improving Unified Large Multimodal Models for Understanding and Generation with Dual Self-Rewards
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- Generalizing vision-language models to novel domains: A comprehensive survey
- AViLA: Asynchronous Vision-Language Agent for Streaming Multimodal Data Interaction
- OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
- RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
- WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
- MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems
- See-in-Pairs: Reference Image-Guided Comparative Vision-Language Models for Medical Diagnosis
- ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
- CLGRPO: Reasoning Ability Enhancement for Small VLMs
- PostAlign: Multimodal Grounding as a Corrective Lens for MLLMs
- SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
- PhysUniBench: A Multi-Modal Physics Reasoning Benchmark at Undergraduate Level
- CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- Taming the Untamed: Graph-Based Knowledge Retrieval and Reasoning for MLLMs to Conquer the Unknown
- Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
- Chiron-o1: Igniting Multimodal Large Language Models towards Generalizable Medical Reasoning via Mentor-Intern Collaborative Search
- Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
- RealSR-R1: Reinforcement Learning for Real-World Image Super-Resolution with Vision-Language Chain-of-Thought
- PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement
- GRPO-CARE: Consistency-Aware Reinforcement Learning for Multimodal Reasoning
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- Can Generalist Vision Language Models (VLMs) Rival Specialist Medical VLMs? Benchmarking and Strategic Insights
- DualTHOR: A Dual-Arm Humanoid Simulation Platform for Contingency-Aware Planning
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-Weights
- IS-Bench: Evaluating Interactive Safety of VLM-Driven Embodied Agents in Daily Household Tasks
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- RiOT: Efficient Prompt Refinement with Residual Optimization Tree
- Demystifying the Visual Quality Paradox in Multimodal Large Language Models
- VLMInferSlow: Evaluating the Efficiency Robustness of Large Vision-Language Models as a Service
- ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving
- LEANN: A Low-Storage Vector Index
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- Understanding GUI Agent Localization Biases through Logit Sharpness
- Show-o2: Improved Native Unified Multimodal Models
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering
- PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
- AGENTSAFE: Benchmarking the Safety of Embodied Agents on Hazardous Instructions
- Exploring MLLMs Perception of Network Visualization Principles
- Play to Generalize: Learning to Reason Through Game Play
- GUI-Robust: A Comprehensive Dataset for Testing GUI Agent Robustness in Real-World Anomalies
- Dense360: Dense Understanding from Omnidirectional Panoramas
- Recognition through Reasoning: Reinforcing Image Geo-localization with Large Vision-Language Models
- LingoLoop Attack: Trapping MLLMs via Linguistic Context and State Entrapment into Endless Loops
- RadFabric: Agentic AI System with Reasoning Capability for Radiology
- Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
- Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning
- X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
- Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
- Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
- Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
- AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
- UltraVideo: High-Quality UHD Video Dataset with Comprehensive Captions
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
- VIS-Shepherd: Constructing Critic for LLM-based Data Visualization Generation
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design
- AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
- Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
- ComplexBench-Edit: Benchmarking Complex Instruction-Driven Image Editing via Compositional Dependencies
- CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making
- Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
- LOP: Learning Optimal Pruning for Efficient On-Demand MLLMs Scaling
- Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation
- MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval
- Benchmarking the Trustworthiness in Multimodal LLMs for Video Understanding
- Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
- EMLoC: Emulator-based Memory-efficient Fine-tuning with LoRA Correction
- DAVID-XR1: Detecting AI-Generated Videos with Explainable Reasoning
- Mitigating Hallucination Through Theory-Consistent Symmetric Multimodal Preference Optimization
- Dynamic Mixture of Curriculum LoRA Experts for Continual Multimodal Instruction Tuning
- Towards Understanding the Cognitive Habits of Large Reasoning Models
- Aligning MLLM Benchmark With Human Preferences via Structural Equation Modeling
- HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
- VGR: Visual Grounded Reasoning
- Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
- Benchmarking Multimodal LLMs on Recognition and Understanding over Chemical Tables
- Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation
- Reasoning Multimodal Large Language Model: Data Contamination and Dynamic Evaluation
- How Important are Videos for Training Video LLMs?
- VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos
- Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification?
- Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs
- Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- Poutine: Vision-Language-Trajectory Pre-Training and Reinforcement Learning Post-Training Enable Robust End-to-End Autonomous Driving
- Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning
- AnimateAnyMesh: A Feed-Forward 4D Foundation Model for Text-Driven Universal Mesh Animation
- LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
- Towards Multimodal Graph Large Language Model
- Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
- Athena: Enhancing Multimodal Reasoning with Data-efficient Process Reward Models
- RoboCerebra: A Large-scale Benchmark for Long-horizon Robotic Manipulation Evaluation
- HeartcareGPT: A Unified Multimodal ECG Suite for Dual Signal-Image Modeling and Understanding
- EASG-Bench: Video Q&A Benchmark with Egocentric Action Scene Graphs
- STSBench: A Spatio-temporal Scenario Benchmark for Multi-modal Large Language Models in Autonomous Driving
- DesignBench: A Comprehensive Benchmark for MLLM-based Front-end Code Generation
- ExAct: A Video-Language Benchmark for Expert Action Analysis
- HMVLM: Multistage Reasoning-Enhanced Vision-Language Model for Long-Tailed Driving Scenarios
- Astra: Toward General-Purpose Mobile Robots via Hierarchical Multimodal Learning
- Movie Facts and Fibs (MF2): A Benchmark for Long Movie Understanding
- VideoMolmo: Spatio-Temporal Grounding Meets Pointing
- Unleashing Hour-Scale Video Training for Long Video-Language Understanding
- AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?
- LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
- MonkeyOCR: Document Parsing with a Structure-Recognition-Relation Triplet Paradigm
- HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning
- ContentV: Efficient Training of Video Generation Models with Limited Compute
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- Look Before You Leap: A GUI-Critic-R1 Model for Pre-Operative Error Diagnosis in GUI Automation
- Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
- Reasoning-Aligned Perception Decoupling for Scalable Multi-modal Reasoning
- VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
- Degradation-Aware Image Enhancement via Vision-Language Classification
- MokA: Multimodal Low-Rank Adaptation for MLLMs
- From Objects to Anywhere: A Holistic Benchmark for Multi-level Visual Grounding in 3D Scenes
- ViDRiP-LLaVA: A Dataset and Benchmark for Diagnostic Reasoning from Pathology Videos
- HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation
- CM1 -- A Dataset for Evaluating Few-Shot Information Extraction with Large Vision Language Models
- How Far Are We from Generating Missing Modalities with Foundation Models?
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- Multimodal Tabular Reasoning with Privileged Structured Information
- ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding
- MiMo-VL Technical Report
- Advancing Multimodal Reasoning: From Optimized Cold Start to Staged Reinforcement Learning
- DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes
- DenseDPO: Fine-Grained Temporal Preference Optimization for Video Diffusion Models
- Seeing What Tastes Good: Revisiting Multimodal Distributional Semantics in the Billion Parameter Era
- Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- Grounded Vision-Language Interpreter for Long-Horizon Bimanual Task and Motion Planning
- Go Beyond Earth: Understanding Human Actions and Scenes in Microgravity Environments
- Rethinking Post-Unlearning Behavior of Large Vision-Language Models
- VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments
- Native-Resolution Image Synthesis
- Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning
- FlySearch: Exploring how vision-language models explore
- Multimodal DeepResearcher: Generating Text-Chart Interleaved Reports From Scratch with Agentic Framework
- SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Seeing the Arrow of Time in Large Multimodal Models
- UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents
- SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
- Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
- Dual-Process Image Generation
- RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
- SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- Hierarchical Intention-Aware Expressive Motion Generation for Humanoid Robots
- MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
- Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
- ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
- SVQA-R1: Reinforcing Spatial Reasoning in MLLMs via View-Consistent Reward Optimization
- MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
- GeoLocSFT: Efficient Visual Geolocation via Supervised Fine-Tuning of Multimodal Foundation Models
- FLEX: A Largescale Multimodal, Multiview Dataset for Learning Structured Representations for Fitness Action Quality Assessment
- ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
- ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding
- Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
- EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
- Fighting Fire with Fire (F3): A Training-free and Efficient Visual Adversarial Example Purification Method in LVLMs
- Affordance Benchmark for MLLMs
- Improve MLLM Benchmark Efficiency through Interview
- GuessBench: Sensemaking Multimodal Creativity in the Wild
- Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
- Ivy-Fake: A Unified Explainable Framework and Benchmark for Image and Video AIGC Detection
- GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking
- Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
- FusionAudio-1.2M: Towards Fine-grained Audio Captioning with Multimodal Contextual Fusion
- Pro3D-Editor : A Progressive-Views Perspective for Consistent and Precise 3D Editing
- Common Inpainted Objects In-N-Out of Context
- MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning
- CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
- Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
- QoQ-Med: Building Multimodal Clinical Foundation Models with Domain-Aware GRPO Training
- RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
- RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
- The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- Reinforcing Video Reasoning with Focused Thinking
- BIMA: Bijective Maximum Likelihood Learning Approach to Hallucination Prediction and Mitigation in Large Vision-Language Models
- SA-Person: Text-Based Person Retrieval with Scene-aware Re-ranking
- MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
- VUDG: A Dataset for Video Understanding Domain Generalization
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning
- SiLVR: A Simple Language-based Video Reasoning Framework
- Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces
- Time Blindness: Why Video-Language Models Can't See What Humans Can?
- MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
- Mixed-R1: Unified Reward Perspective For Reasoning Capability in Multimodal Large Language Models
- Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
- GenSpace: Benchmarking Spatially-Aware Image Generation
- InterMT: Multi-Turn Interleaved Preference Alignment with Human Feedback
- ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
- Grounded Reinforcement Learning for Visual Reasoning
- Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information
- OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
- VAU-R1: Advancing Video Anomaly Understanding via Reinforcement Fine-Tuning
- VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
- Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- InfiMed: Low-Resource Medical MLLMs with Advancing Understanding and Reasoning
- Generalizable Video Quality Assessment via Weak-to-Strong Learning
- SNS-Bench-VL: Benchmarking Multimodal Large Language Models in Social Networking Services
- Synthetic Document Question Answering in Hungarian
- LLM-based HSE Compliance Assessment: Benchmark, Performance, and Advancements
- To Trust Or Not To Trust Your Vision-Language Model's Prediction
- Elicit and Enhance: Advancing Multimodal Reasoning in Medical Scenarios
- Are MLMs Trapped in the Visual Room?
- PixelThink: Towards Efficient Chain-of-Pixel Reasoning
- VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC Videos
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- EndoBench: A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis
- Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
- Generating Fit Check Videos with a Handheld Camera
- TrackVLA: Embodied Visual Tracking in the Wild
- Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
- DIP-R1: Deep Inspection and Perception with RL Looking Through and Understanding Complex Scenes
- OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
- VidText: Towards Comprehensive Evaluation for Video Text Understanding
- Scaling-up Perceptual Video Quality Assessment
- EvoMoE: Expert Evolution in Mixture of Experts for Multimodal Large Language Models
- Reinforced Reasoning for Embodied Planning
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
- Fostering Video Reasoning via Next-Event Prediction
- Balanced Token Pruning: Accelerating Vision Language Models Beyond Local Optimization
- EF-VI: Enhancing End-Frame Injection for Video Inbetweening
- Sherlock: Self-Correcting Reasoning in Vision-Language Models
- VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
- VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
- Universal Visuo-Tactile Video Understanding for Embodied Interaction
- SAM-R1: Leveraging SAM for Reward Feedback in Multimodal Segmentation via Reinforcement Learning
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
- AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation
- Advancing Multimodal Reasoning via Reinforcement Learning with Cold Start
- UI-Evol: Automatic Knowledge Evolving for Computer Use Agents
- Pearl: A Multimodal Culturally-Aware Arabic Instruction Dataset
- A2Seek: Towards Reasoning-Centric Benchmark for Aerial Anomaly Understanding
- DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion Model
- SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model
- ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- MSEarth: A Multimodal Scientific Dataset and Benchmark for Phenomena Uncovering in Earth Science
- MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
- PIPE: Physics-Informed Position Encoding for Alignment of Satellite Images and Time Series
- Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs
- ViewSpatial-Bench: Evaluating Multi-perspective Spatial Localization in Vision-Language Models
- TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios
- Roboflow100-VL: A Multi-Domain Object Detection Benchmark for Vision-Language Models
- Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
- DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
- Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
- DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
- A Stereotype Content Analysis on Color-related Social Bias in Large Vision Language Models
- DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
- HoliTom: Holistic Token Merging for Fast Video Large Language Models
- UI-Genie: A Self-Improving Approach for Iteratively Boosting MLLM-based Mobile GUI Agents
- RefAV: Towards Planning-Centric Scenario Mining
- IndustryEQA: Pushing the Frontiers of Embodied Question Answering in Industrial Scenarios
- Evaluating and Steering Modality Preferences in Multimodal Large Language Model
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
- Cross from Left to Right Brain: Adaptive Text Dreamer for Vision-and-Language Navigation
- HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
- CPathAgent: An Agent-based Foundation Model for Interpretable High-Resolution Pathology Image Analysis Mimicking Pathologists' Diagnostic Logic
- What Changed? Detecting and Evaluating Instruction-Guided Image Edits with Multimodal Large Language Models
- MangaVQA and MangaLMM: A Benchmark and Specialized Model for Multimodal Manga Understanding
- OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- MMGeoLM: Hard Negative Contrastive Learning for Fine-Grained Geometric Understanding in Large Multimodal Models
- MineAnyBuild: Benchmarking Spatial Planning for Open-world AI Agents
- Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought
- Two Causally Related Needles in a Video Haystack
- USB: A Comprehensive and Unified Safety Evaluation Benchmark for Multimodal Large Language Models
- Align and Surpass Human Camouflaged Perception: Visual Refocus Reinforcement Fine-Tuning
- JailBound: Jailbreaking Internal Safety Boundaries of Vision-Language Models
- Guard Me If You Know Me: Protecting Specific Face-Identity from Deepfakes
- What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Enhancing Visual Reliance in Text Generation: A Bayesian Perspective on Mitigating Hallucination in Large Vision-Language Models
- VLMLight: Safety-Critical Traffic Signal Control via Vision-Language Meta-Control and Dual-Branch Reasoning Architecture
- Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
- VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models
- Thinking with Visual Abstract: Enhancing Multimodal Reasoning via Visual Abstraction
- ReaMOT: A Benchmark and Framework for Reasoning-based Multi-Object Tracking
- MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval
- MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models
- VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
- ImgEdit: A Unified Image Editing Dataset and Benchmark
- Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
- Benchmarking Multimodal Knowledge Conflict for Large Multimodal Models
- Universal Reasoner: A Single, Composable Plug-and-Play Reasoner for Frozen LLMs
- Increasing Computation Resolves Conflicts in Vision Language Models
- Shifting AI Efficiency From Model-Centric to Data-Centric Compression
- SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring
- RTime-QA: A Benchmark for Atomic Temporal Event Understanding in Large Multi-modal Models
- VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
- InfoChartQA: A Benchmark for Multimodal Question Answering on Infographic Charts
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models
- SeePhys: Does Seeing Help Thinking? -- Benchmarking Vision-Based Physics Reasoning
- Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning
- SD-OVON: A Semantics-aware Dataset and Benchmark Generation Pipeline for Open-Vocabulary Object Navigation in Dynamic Scenes
- ToDRE: Effective Visual Token Pruning via Token Diversity and Task Relevance
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- So-Fake: Benchmarking and Explaining Social Media Image Forgery Detection
- MLLMs are Deeply Affected by Modality Bias
- ThanoRA: Task Heterogeneity-Aware Multi-Task Low-Rank Adaptation
- Flex-Judge: Text-Only Reasoning Unleashes Zero-Shot Multimodal Evaluators
- Rethinking Causal Mask Attention for Vision-Language Inference
- CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- Generative RLHF-V: Learning Principles from Multi-modal Human Preference
- ReasonMap: Towards Fine-Grained Visual Reasoning from Transit Maps
- Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit
- DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and Understanding
- Seeing Beyond Words: MatVQA for Challenging Visual-Scientific Reasoning in Materials Science
- Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence
- One RL to See Them All: Visual Triple Unified Reinforcement Learning
- Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
- AVerImaTeC: A Dataset for Automatic Verification of Image-Text Claims with Evidence from the Web
- Douyin Multimodal Embedding Model Technical Report
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
- Decoupling semantics from vision: A framework for faithful visual-text compression evaluation
- Token Reduction Should Go Beyond Efficiency in Generative Models -- From Vision, Language to Multimodality
- Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models
- Towards General Continuous Memory for Vision-Language Models
- InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
- Illuminating Visual Identity in Universal Multimodal Embeddings
- Co-Reinforcement Learning for Unified Multimodal Understanding and Generation
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- MemeReaCon: Probing Contextual Meme Understanding in Large Vision-Language Models
- VIBE: Annotation-Free Video-to-Text Information Bottleneck Evaluation for TL;DR
- FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow
- CHAOS: Chart Analysis with Outlier Samples
- Let Androids Dream of Electric Sheep: A Human-Inspired Image Implication Understanding and Reasoning Framework
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
- OpenSeg-R: Improving Open-Vocabulary Segmentation via Step-by-Step Visual Reasoning
- MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning
- Think or Not? Selective Reasoning via Reinforcement Learning for Vision-Language Models
- V2V: Scaling Event-Based Vision through Efficient Video-to-Voxel Simulation
- OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs
- SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models
- Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
- From Evaluation to Defense: Advancing Safety in Video Large Language Models
- Mitigating Hallucinations in Vision-Language Models through Image-Guided Head Suppression
- NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
- SEM: Enhancing Spatial Understanding for Robust Robot Manipulation
- Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning
- GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- Circle-RoPE: Cone-like Decoupled Rotary Positional Embedding for Large Vision-Language Models
- Implicit Jailbreak Attacks via Cross-Modal Information Concealment on Vision-Language Models
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
- Recursive Offloading for LLM Serving in Multi-tier Networks
- STAR-R1: Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- IA-T2I: Internet-Augmented Text-to-Image Generation
- NeSyGeo: A Neuro-Symbolic Framework for Multimodal Geometric Reasoning Data Generation
- Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
- ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
- Better Safe Than Sorry? Overreaction Problem of Vision Language Models in Visual Emergency Recognition
- P2P: Automated Paper-to-Poster Generation and Fine-Grained Benchmark
- AvatarShield: Visual Reinforcement Learning for Human-Centric Synthetic Video Detection
- TACO: Enhancing Multimodal In-context Learning via Task Mapping-Guided Sequence Configuration
- GUI-G1: Understanding R1-Zero-Like Training for Visual Grounding in GUI Agents
- InstructSAM: A Training-Free Framework for Instruction-Oriented Remote Sensing Object Recognition
- Web-Shepherd: Advancing PRMs for Reinforcing Web Agents
- LENS: Multi-level Evaluation of Multimodal Reasoning with Large Language Models
- Can VLMs Detect and Localize Fine-Grained AI-Edited Images?
- Dynamic Resolution Routing for Efficient Egocentric Grounding
- Entailed Opinion Matters: Improving the Fact-Checking Performance of Language Models by Relying on their Entailment Ability
- Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
- ReGUIDE: Data Efficient GUI Grounding via Spatial Reasoning and Search
- MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
- Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
- Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
- LiveVLM: Efficient Online Video Understanding via Streaming-Oriented KV Cache and Retrieval
- Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models
- Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
- Multimodal Cultural Safety: Evaluation Framework and Alignment Strategies
- EmoSign: A Multimodal Dataset for Understanding Emotions in American Sign Language
- Emerging Properties in Unified Multimodal Pretraining
- ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions
- VoQA: Visual-only Question Answering
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
- Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
- Visual Agentic Reinforcement Fine-Tuning
- LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts
- Visual Instruction Bottleneck Tuning
- VisualQuality-R1: Reasoning-Induced Image Quality Assessment via Reinforcement Learning to Rank
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
- VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- Memory-Centric Embodied Question Answering
- Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting
- Scalable Bayesian Monte Carlo: fast uncertainty estimation beyond deep ensembles
- Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
- FLASH: Latent-Aware Semi-Autoregressive Speculative Decoding for Multimodal Tasks
- BusterX: MLLM-Powered AI-Generated Video Forgery Detection and Explanation
- GEM: Gaussian Embedding Modeling for Out-of-Distribution Detection in GUI Agents
- HydraInfer: Hybrid Disaggregated Scheduling for Multimodal Large Language Model Serving
- Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
- G1: Bootstrapping Perception and Reasoning Abilities of Vision-Language Model via Reinforcement Learning
- SounDiT: Geo-Contextual Soundscape-to-Landscape Generation
- DreamGen: Unlocking Generalization in Robot Learning through Video World Models
- AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
- SurveillanceVQA-589K: A Benchmark for Comprehensive Surveillance Video-Language Understanding with Large Models
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
- MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
- Incentivizing Multimodal Reasoning in Large Models for Direct Robot Manipulation
- RoboFAC: A Comprehensive Framework for Robotic Failure Analysis and Correction
- Enhancing Visual Grounding for GUI Agents via Self-Evolutionary Reinforcement Learning
- GUI-Shift: Enhancing VLM-Based GUI Agents through Self-supervised Reinforcement Learning
- Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
- Visuospatial Cognitive Assistant
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- VideoRFT: Incentivizing Video Reasoning Capability in MLLMs via Reinforced Fine-Tuning
- CorBenchX: Large-Scale Chest X-Ray Error Dataset and Vision-Language Model Benchmark for Report Error Correction
- SafeVid: Toward Safety Aligned Video Large Multimodal Models
- GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
- Remember-R1: Mitigating Long-Context Visual Forgetting through Reinforcement Learning
- Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language Navigation
- LOVE: Benchmarking and Evaluating Text-to-Video Generation and Video-to-Text Interpretation
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Are vision language models robust to uncertain inputs?
- MedSG-Bench: A Benchmark for Medical Image Sequences Grounding
- Location-Aware Fine-Grained Representation Learning for Medical Vision Foundation Models
- PSDiffusion: Harmonized Multi-Layer Image Generation via Layout and Appearance Alignment
- Visual Planning: Let's Think Only with Images
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- InfantAgent-Next: A Multimodal Generalist Agent for Automated Computer Interaction
- TAIJI: MCP-based Multi-Modal Data Analytics on Data Lakes
- Creating General User Models from Computer Use
- Group-in-Group Policy Optimization for LLM Agent Training
- SpatialAfford: Teaching Compact VLMs Where to Look and Where to Ground for Affordance
- HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- VISTA: Enhancing Vision-Text Alignment in MLLMs via Cross-Modal Mutual Information Maximization
- WebInject: Prompt Injection Attack to Web Agents
- GeoGrid-Bench: Can Foundation Models Understand Multimodal Gridded Geo-Spatial Data?
- Cross-Image Contrastive Decoding: Precise, Lossless Suppression of Language Priors in Large Vision-Language Models
- DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards
- Why 1 + 1 < 1 in Visual Token Pruning: Beyond Naive Integration via Multi-Objective Balanced Covering
- PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language
- MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models
- PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
- Task-Core Memory Management and Consolidation for Long-term Continual Learning
- VTLA: Vision-Tactile-Language-Action Model with Preference Learning for Insertion Manipulation
- Emotion Knowledge Enhancement for Vision Large Language Models: A Self-Verification Approach for High-Quality Emotion Instruction Data Generation
- Bridging Human Oversight and Black-box Driver Assistance: Vision-Language Models for Predictive Alerting in Lane Keeping Assist Systems
- ManipBench: Benchmarking Vision-Language Models for Low-Level Robot Manipulation
- Qwen3 Technical Report
- Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving
- SafeBuild-Bench: A Temporal-Robust Construction Safety Benchmark with Graph-Enhanced Data Mining
- Extracting Visual Plans from Unlabeled Videos via Symbolic Guidance
- Decoding Neighborhood Environments with Large Language Models
- DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams
- DeceptionX: From Multimodal Evidence to Explainable Deception Detection
- Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
- Large Language Models for Computer-Aided Design: A Survey
- Relative Overfitting and Accept-Reject Framework
- Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning
- Ophora: A Large-Scale Data-Driven Text-Guided Ophthalmic Surgical Video Generation Model
- DriveSOTIF: Advancing Perception SOTIF Through Multimodal Large Language Models
- MELLM: Exploring LLM-Powered Micro-Expression Understanding Enhanced by Subtle Motion Perception
- Bridging Ears and Eyes: Analyzing Audio and Visual Large Language Models to Humans in Visible Sound Recognition and Reducing Their Sensory Gap via Cross-Modal Distillation
- DriveCode: Domain Specific Numerical Encoding for LLM-Based Autonomous Driving
- V2VCrafter: Consistent Street-View Image Generation Across Vehicles
- Describe Anything in Medical Images
- SITE: towards Spatial Intelligence Thorough Evaluation
- DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning
- ARGen: Affect-Reinforced Generative Augmentation towards Vision-based Dynamic Emotion Perception
- Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
- Learning to Select Visual In-Context Demonstrations
- SVRepair: Structured Visual Reasoning for Automated Program Repair
- Chart Specification: Structural Representations for Incentivizing VLM Reasoning in Chart-to-Code Generation
- NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
- Retrieval-augmented in-context learning for multimodal large language models in disease classification
- Interleave-VLA: Enhancing Robot Manipulation with Interleaved Image-Text Instructions
- DIAL: Decoupling Intent and Action via Latent World Modeling for End-to-End VLA
- Transferable Adversarial Attacks on Black-Box Vision-Language Models
- VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
- Controllable Weather Synthesis and Removal with Video Diffusion Models
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
- Visual Test-time Scaling for GUI Agent Grounding
- MINERVA: Evaluating Complex Video Reasoning
- A Survey of OCR Evaluation Methods and Metrics and the Invisibility of Historical Documents
- Kwai Keye-VL-2.0 Technical Report
- SpaceMind++: Toward Allocentric Cognitive Maps for Spatially Grounded Video MLLMs
- The Price Is Not Right: Neuro-Symbolic Methods Outperform VLAs on Structured Long-Horizon Manipulation Tasks with Significantly Lower Energy Consumption
- FlashPrefill: Instantaneous Pattern Discovery and Thresholding for Ultra-Fast Long-Context Prefilling
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing
- Better, Stronger, Faster, and Broader: Structured All-Mask Prediction for MLLM-Based Segmentation
- Beyond the Single Camera: Agentic Multi-View Reasoning in Sports Video Understanding
- Infinity-Parser2 Technical Report
- FINER: MLLMs Hallucinate under Fine-grained Negative Queries
- From Fixed to Free Cameras: Calibration-Free View-Robust Vision-Language-Action Model
- SIFT: Self-Imagination Fine-Tuning for Physically Plausible Motion in Video Diffusion Models
- Anticipatory Planning for Multimodal AI Agents
- V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning
- ScaleTrack: Scaling and back-tracking Automated GUI Agents
- Think-as-You-See: Streaming Chain-of-Thought Reasoning for Large Vision-Language Models
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
- InterleaveThinker: Reinforcing Agentic Interleaved Generation
- HECTOR: Hybrid Editable Compositional Object References for Video Generation
- Lightweight Visual Reasoning for Socially-Aware Robots
- FireRed-OCR Technical Report
- Towards Efficient Online Tuning of VLM Agents via Counterfactual Soft Reinforcement Learning
- AVA: Towards Agentic Video Analytics with Vision Language Models
- VLANeXt: Recipes for Building Strong VLA Models
- Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- AI-Model Network: Concept, Current State and Future
- AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human Images
- Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
- CMT: A Cascade MAR with Topology Predictor for Multimodal Conditional CAD Generation
- In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- DEEMO: De-identity Multimodal Emotion Recognition and Reasoning
- MEG-RAG: Quantifying Multi-modal Evidence Grounding for Evidence Selection in RAG
- Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
- ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
- Learning Action Priors for Cross-embodiment Robot Manipulation
- PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training
- OmniVerifier-M1: Multimodal Meta-Verifier with Explicit Structured Recalibration
- Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
- World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
- LLaDA2.0-Uni: Unifying Multimodal Understanding and Generation with Diffusion Large Language Model
- Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
- Predicting Camera Pose from Perspective Descriptions for Spatial Reasoning
- RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks
- Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
- MLDocRAG: Multimodal Long-Context Document Retrieval Augmented Generation
- SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL
- SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking
- Fast-Slow Thinking GRPO for Large Vision-Language Model Reasoning
- Revisiting Data Auditing in Large Vision-Language Models
- Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
- Mosaic: Cross-Modal Clustering for Efficient Video Understanding
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- Okara: Detection and Attribution of TLS Man-in-the-Middle Vulnerabilities in Android Apps with Foundation Models
- DataCross: A Unified Benchmark and Agent Framework for Cross-Modal Heterogeneous Data Analysis
- SLAP: Scalable Language-Audio Pretraining with Variable-Duration Audio and Multi-Objective Training
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- GSTEP: Global Spatio-Temporal Density-Driven Visual Token Pruning for Efficient Video Large Language Models
- Adaptive Two-Stage Visual Token Pruning for Efficient Inference in Video-Language Models
- VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
- OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet
- When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
- UHP Detection: LVLMs have their Unique Hallucination Pattern in the Consistency Space
- How to Discover Knowledge for FutureG: Contextual RAG and LLM Prompting for O-RAN
- Beyond Clicking:A Step Towards Generalist GUI Grounding via Text Dragging
- Step1X-Edit: A Practical Framework for General Image Editing
- DIMT25@ICDAR2025: HW-TSC's End-to-End Document Image Machine Translation System Leveraging Large Vision-Language Model
- TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary Generation
- MASR: Self-Reflective Reasoning through Multimodal Hierarchical Attention Focusing for Agent-based Video Understanding
- Benchmarking Multimodal Mathematical Reasoning with Explicit Visual Dependency
- DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
- BadVideo: Stealthy Backdoor Attack against Text-to-Video Generation
- Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
- URECA: Unique Region Caption Anything
- Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
- VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
- Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning
- From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
- Vidi: Large Multimodal Models for Video Understanding and Editing
- FaceInsight: A Multimodal Large Language Model for Face Perception
- MMInference: Accelerating Pre-filling for Long-Context VLMs via Modality-Aware Permutation Sparse Attention
- TrustGeoGen: Formal-Verified Data Engine for Trustworthy Multi-modal Geometric Problem Solving
- VISTA-OCR: Towards generative and interactive end to end OCR models
- MR. Video: "MapReduce" is the Principle for Long Video Understanding
- LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
- Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform
- IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- Towards Understanding Camera Motions in Any Video
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding
- M2IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
- Relation-R1: Progressively Cognitive Chain-of-Thought Guided Reinforcement Learning for Unified Relation Comprehension
- How Well Can General Vision-Language Models Learn Medicine By Watching Public Educational Videos?
- InfiGUI-R1: Advancing Multimodal GUI Agents from Reactive Actors to Deliberative Reasoners
- Learning Joint ID-Textual Representation for ID-Preserving Image Synthesis
- VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
- Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
- Chain-of-Evidence Multimodal Reasoning for Few-shot Temporal Action Localization
- Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
- VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization
- Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning
- NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
- Enhancing the Geometric Problem-Solving Ability of Multimodal LLMs via Symbolic-Neural Integration
- LAD-Reasoner: Tiny Multimodal Models are Good Reasoners for Logical Anomaly Detection
- Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
- Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
- AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
- FocusedAD: Character-centric Movie Audio Description
- Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
- Beyond Relevance: Bayesian Evidence Acquisition for Agentic Whole-Slide Image Reasoning
- Beyond Frame Selection: Rethinking Long-Video Understanding with MLLMs
- TruthLens: Object Hallucination Detection via Self-Evaluating Truthfulness Scores in LVLMs
- OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction
- Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?
- M3Prune: Hierarchical Collaborative Pruning for Efficient Multi-Modal Multi-Agent Retrieval-Augmented Generation
- Visual Grounding in Zero-Shot Vision-Language Control
- VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances
- HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
- SOVABench: A Vehicle Surveillance Action Retrieval Benchmark for Multimodal Large Language Models
- Consensus Entropy: Harnessing Multi-VLM Agreement for Self-Verifying and Self-Improving OCR
- MMC: Iterative Refinement of VLM Reasoning via MCTS-based Multimodal Critique
- HippoMM: Hippocampal-inspired Multimodal Memory for Long Audiovisual Event Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- FingER: Content Aware Fine-grained Evaluation with Reasoning for AI-Generated Videos
- Mavors: Multi-granularity Video Representation for Multimodal Large Language Model
- GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
- HistLLM: A Unified Framework for LLM-Based Multimodal Recommendation with User History Encoding and Compression
- NTIRE 2025 Challenge on Cross-Domain Few-Shot Object Detection: Methods and Results
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- GeoNav: Empowering MLLMs with Explicit Geospatial Reasoning Abilities for Language-Goal Aerial Navigation
- FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality Assessment
- VideoAds for Fast-Paced Video Understanding
- PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
- Towards Explainable Partial-AIGC Image Quality Assessment
- LMM4LMM: Benchmarking and Evaluating Large-multimodal Image Generation with LMMs
- VCR-Bench: A Comprehensive Evaluation Framework for Video Chain-of-Thought Reasoning
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
- TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs
- FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
- Kimi-VL Technical Report
- VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning
- OmniCaptioner: One Captioner to Rule Them All
- Distilling Textual Priors from LLM to Efficient Image Fusion
- V-MAGE: A Game Evaluation Framework for Assessing Vision-Centric Capabilities in Multimodal Large Language Models
- MDK12-Bench: A Multi-Discipline Benchmark for Evaluating Reasoning in Multimodal Large Language Models
- On the Suitability of Reinforcement Fine-Tuning to Visual Tasks
- Mind the (Data) Gap: Evaluating Vision Systems in Small Data Applications
- Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought
- Don't Lag, RAG: Training-Free Adversarial Detection Using RAG
Discussions
Related