Sigmoid Loss for Language Image Pre-Training
2023/03/27 by Xiaohua Zhai, Zhai, Xiaohua, Basil Mustafa +6 · 3 voices · 717 citations
Computer Science · #Advanced Neural Network Applications #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications #cs.AI #cs.CV
paper · pdf · doi:10.48550/arxiv.2303.15343
openalex publication_date 2023/03/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
We propose a simple pairwise Sigmoid loss for Language-Image Pre-training (SigLIP). Unlike standard contrastive learning with softmax normalization, the sigmoid loss operates solely on image-text pairs and does not require a global view of the pairwise similarities for normalization. The sigmoid loss simultaneously allows further scaling up the batch size, while also performing better at smaller batch sizes. Combined with Locked-image Tuning, with only four TPUv4 chips, we train a SigLiT model that achieves 84.5% ImageNet zero-shot accuracy in two days. The disentanglement of the batch size from the loss further allows us to study the impact of examples vs pairs and negative to positive ratio. Finally, we push the batch size to the extreme, up to one million, and find that the benefits of growing batch size quickly diminish, with a more reasonable batch size of 32k being sufficient. We release our models at https://github.com/google-research/bigvision and hope our research motivates further explorations in improving the quality and efficiency of language-image pre-training.
Cited by
- Scaling Native Multimodal Pre-Training From Scratch
- INQUIRE-Search: Interactive Discovery in Large-Scale Biodiversity Databases
- Anticipate Before Acting: Future-State-Conditioned Vision-Language Navigation
- MultiRef-Compass: Towards Comprehensive Evaluation of Multi-Reference-to-Audio-Video Generation
- Art Beyond Semantics: Sheaf-Informed Contrastive Learning for Multi-Relational Representations
- LaViDa: A Large Diffusion Language Model for Multimodal Understanding
- Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
- DINOde: Continuous Vision-Text Alignment for Open-Vocabulary Semantic Segmentation
- ReferTrack: Referring Then Tracking for Embodied Visual Tracking
- Generative Semantic Multi-Object Tracking: A Large-Scale Benchmark and an MLLM-Driven Reasoning Framework
- Pixel-Space Diffusion Transformers
- Mitigating Modality and Language-Style Gaps for Zero-Shot Video Moment Retrieval
- LatentLens: Revealing Highly Interpretable Visual Tokens in LLMs
- Don't Fool Me Twice: Adapting to Adversity in the Wild with Experience-Driven Reasoning
- Evaluating Uncertainty and Quality of Visual Language Action-enabled Robots
- GeoTrace: Geometry-Aware Trajectory Token Compression for Video Large Language Models
- Beyond Objective Expressivity: Geometry Preservation in Multimodal Contrastive Learning
- D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models
- Semantic Richness or Geometric Reasoning? The Fragility of VLM's Visual Invariance
- DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
- The JEPA Predictor: A Transferable Operator for Occluded Feature Completion
- Parameter-Efficient Adaptation of a Multi-Stream Vision-Language Framework for Blind Image Quality Assessment
- The Many Senses of Visual Similarity: A Text-Prompted Image Perceptual Metric
- Patch Policy: Efficient Embodied Control via Dense Visual Representations
- FM-VLA: Force-based Memory for Vision-Language-Action Models in Contact-Rich Manipulation
- Three-Body Scattering for Generative Modeling
- Enhancing Vision Foundation Models via Multimodal Continual Pre-Training
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local Attenders
- HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
- REGLUE Your Latents with Global and Local Semantics for Entangled Diffusion
- Learning to Fold: prizewinning solution at LeHome Challenge 2026 (1st place online, 2nd offline)
- RhinoVLA Technical Report
- Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- SceneBind: Binding What and Where Across Vision, Audio and Language
- A Modern Multimodal Assistant on a 6 GB 2011 GPU: Stage-Validated, All-GPU CUDA Inference for Fermi
- Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding
- Trajectory-aware Cross-view Geo-localization with Sequential Observations
- Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings
- 3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism
- BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges
- Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA
- Cross-Modal Taxonomic Generalization in (Vision-) Language Models
- LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
- Cameras as Relative Positional Encoding
- Token Bottleneck: One Token to Remember Dynamics
- WorldVLA: Towards Autoregressive Action World Model
- PanSt3R: Multi-view Consistent Panoptic Segmentation
- JAFAR: Jack up Any Feature at Any Resolution
- QuARI: Query Adaptive Retrieval Improvement
- Perception Encoder: The best visual embeddings are not at the output of the network
- Interpreting the linear structure of vision-language model embedding spaces
- Visual Language Models show widespread visual deficits on neuropsychological tests
- Transfer between Modalities with MetaQueries
- RANa: Retrieval-Augmented Navigation
- Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
- Vision as LoRA
- OpenCity3D: What do Vision-Language Models know about Urban Environments?
- ILIAS: Instance-Level Image retrieval At Scale
- TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
- Rethinking the Use of Vision Transformers for AI-Generated Image Detection
- Distribution Matching Variational AutoEncoder
- ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- Act2Goal: From World Model To General Goal-conditioned Policy
- VGGT-Ω
- Utonia: Toward One Encoder for All Point Clouds
- Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
- Retrieve and Segment: Are a Few Examples Enough to Bridge the Supervision Gap in Open-Vocabulary Segmentation?
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- VLM-PAR: A Vision Language Model for Pedestrian Attribute Recognition
- Open-Source Multimodal Moxin Models with Moxin-VLM and Moxin-VLA
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- MAViE: A Multi-scale Adaptive Vision Encoder for Fine-grained Visual Perception and Efficient Multimodal Reasoning
- What CLIP Knows but Cannot Say: Recovering Negation from Frozen Intermediate Features
- The Spatial Blindspot of Vision-Language Models
- Similarity Is Not Logic: Factored Inference for Dual-Encoder Vision-Language Models
- Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation
- IMPRINT: Image-Conditioned Query Enrichment for Long-Tail Object Goal Navigation
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- Controlling Embedding Spaces with Text-Conditioned Transformations
- RoboMME-Interference: Benchmarking Robot Memory Under Interference
- Unlocking Spatial Grounding in Large Audio-Visual Retrieval models
- Lexical discovery in unknown environments orchestrated by Large Language Models
- CHaystack: Benchmarking Chinese Document Retrieval and VQA
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Context Sensitivity Improves Human-Machine Visual Alignment
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- RLLaVA: An RL-central Framework for Language and Vision Assistants
- Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Learning from Next-Frame Prediction: Autoregressive Video Modeling Encodes Effective Representations
- Towards Natural Language-Based Document Image Retrieval: New Dataset and Benchmark
- Detecting Non-Optimal Decisions of Embodied Agents via Diversity-Guided Metamorphic Testing
- Beyond Vision: Contextually Enriched Image Captioning with Multi-Modal Retrieval
- SigLino: Efficient Multi-Teacher Distillation for Agglomerative Vision Foundation Models
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- IPCV: Information-Preserving Compression for MLLM Visual Encoders
- Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models
- LLaViDA: A Large Language Vision Driving Assistant for Explicit Reasoning and Enhanced Trajectory Planning
- Investigating Spatial Attention Bias in Vision-Language Models
- Faithful and Stable Neuron Explanations for Trustworthy Mechanistic Interpretability
- Name That Part: 3D Part Segmentation and Naming
- MMLANDMARKS: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering
- Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
- ABE-CLIP: Training-Free Attribute Binding Enhancement for Compositional Image-Text Matching
- RoomEditor++: A Parameter-Sharing Diffusion Architecture for High-Fidelity Furniture Synthesis
- The Effect of Negation on CLIP in Medical Imaging: Limitations of Contrastive Language-Image Pretraining
- 4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- A Benchmark and Agentic Framework for Omni-Modal Reasoning and Tool Use in Long Videos
- Collaborative Edge-to-Server Inference for Vision-Language Models
- Sceniris: A Fast Procedural Scene Generation Framework
- Vision-Language-Action Models for Autonomous Driving: Past, Present, and Future
- Are vision-language models ready to zero-shot replace supervised classification models in agriculture?
- In Pursuit of Pixel Supervision for Visual Pre-training
- MoonSeg3R: Monocular Online Zero-Shot Segment Anything in 3D with Reconstructive Foundation Priors
- MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
- Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- Unified Semantic Transformer for 3D Scene Understanding
- EXAONE Path 2.5: Pathology Foundation Model with Multi-Omics Alignment
- SuperCLIP: CLIP with Simple Classification Supervision
- Directional Textual Inversion for Personalized Text-to-Image Generation
- SocialNav-MoE: A Mixture-of-Experts Vision Language Model for Socially Compliant Navigation with Reinforcement Fine-Tuning
- Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- MineTheGap: Automatic Mining of Biases in Text-to-Image Models
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
- Patch-wise Retrieval: A Bag of Practical Techniques for Instance-level Matching
- V-Rex: Real-Time Streaming Video LLM Acceleration via Dynamic KV Cache Retrieval
- Moment and Highlight Detection via MLLM Frame Segmentation
- RePack then Refine: Efficient Diffusion Transformer with Vision Foundation Model
- Reconstruction as a Bridge for Event-Based Visual Question Answering
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- An Anatomy of Vision-Language-Action Models: From Modules to Milestones and Challenges
- VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing
- UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
- Image Tiling for High-Resolution Reasoning: Balancing Local Detail with Global Context
- Learning complete and explainable visual representations from itemized text supervision
- Vision-Language Models for Infrared Industrial Sensing in Additive Manufacturing Scene Description
- VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- DuetSVG: Unified Multimodal SVG Generation with Internal Visual Guidance
- SoccerMaster: A Vision Foundation Model for Soccer Understanding
- Interpretable and Steerable Concept Bottleneck Sparse Autoencoders
- Towards Accessible Physical AI: LoRA-Based Fine-Tuning of VLA Models for Real-World Robot Control
- DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
- HiF-VLA: Hindsight, Insight and Foresight through Motion Representation for Vision-Language-Action Models
- GLaD: Geometric Latent Distillation for Vision-Language-Action Models
- SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
- Towards Lossless Ultimate Vision Token Compression for VLMs
- Beyond the Noise: Aligning Prompts with Latent Representations in Diffusion Models
- The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Bridging Scale Discrepancies in Robotic Control via Language-Based Action Representations
- Aerial Vision-Language Navigation with a Unified Framework for Spatial, Temporal and Embodied Reasoning
- Relational Visual Similarity
- OpenVE-3M: A Large-Scale High-Quality Dataset for Instruction-Guided Video Editing
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video Understanding
- All You Need Are Random Visual Tokens? Demystifying Token Pruning in VLLMs
- A Geometric Unification of Concept Learning with Concept Cones
- Generalized Referring Expression Segmentation on Aerial Photos
- MulCLIP: A Multi-level Alignment Framework for Enhancing Fine-grained Long-context CLIP
- Task adaptation of Vision-Language-Action model: 1st Place Solution for the 2025 BEHAVIOR Challenge
- Less Is More, but Where? Dynamic Token Compression via LLM-Guided Keyframe Prior
- VideoVLA: Video Generators Can Be Generalizable Robot Manipulators
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies
- Mitigating Bias with Words: Inducing Demographic Ambiguity in Face Recognition Templates by Text Encoding
- EmoStyle: Emotion-Driven Image Stylization
- Intra-Class Probabilistic Embeddings for Uncertainty Estimation in Vision-Language Models
- EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
- The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems
- Not All Birds Look The Same: Identity-Preserving Generation For Birds
- SP-Det: Self-Prompted Dual-Text Fusion for Generalized Multi-Label Lesion Detection
- PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
- VAT: Vision Action Transformer by Unlocking Full Representation of ViT
- Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation
- Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling
- Hierarchical Process Reward Models are Symbolic Vision Learners
- Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models
- MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- Learning What to Attend First: Modality-Importance-Guided Reasoning for Reliable Multimodal Emotion Understanding
- WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
- VLM-Pruner: Buffering for Spatial Sparsity in an Efficient VLM Centrifugal Token Pruning Paradigm
- TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
- ManualVLA: A Unified VLA Model for Chain-of-Thought Manual Generation and Robotic Manipulation
- Learning Sim-to-Real Humanoid Locomotion in 15 Minutes
- Script: Graph-Structured and Query-Conditioned Semantic Token Pruning for Multimodal Large Language Models
- DiG-Flow: Discrepancy-Guided Flow Matching for Robust VLA Models
- InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
- TokenPure: Watermark Removal through Tokenized Appearance and Structural Guidance
- Assimilation Matters: Model-level Backdoor Detection in Vision-Language Pretrained Models
- Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
- CycliST: A Video Language Model Benchmark for Reasoning on Cyclical State Transitions
- TRoVe: Discovering Error-Inducing Static Feature Biases in Temporal Vision-Language Models
- SwiftVLA: Unlocking Spatiotemporal Dynamics for Lightweight VLA Models at Minimal Overhead
- HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Optimizing LVLMs with On-Policy Data for Effective Hallucination Mitigation
- Describe Anything Anywhere At Any Moment
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction
- Optimizing Multimodal Language Models through Attention-based Interpretability
- PowerCLIP: Powerset Alignment for Contrastive Pre-Training
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language Models
- GSPN-2: Efficient Parallel Sequence Modeling
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- Attention-Guided Patch-Wise Sparse Adversarial Attacks on Vision-Language-Action Models
- E0: Enhancing Generalization and Fine-Grained Control in VLA Models via Continuized Discrete Diffusion
- uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
- CameraMaster: Unified Camera Semantic-Parameter Control for Photography Retouching
- FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning
- G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
- When Robots Obey the Patch: Universal Transferable Patch Attacks on Vision-Language-Action Models
- RLM: A Vision-Language Model Approach for Radar Scene Understanding
- BotaCLIP: Contrastive Learning for Botany-Aware Representation of Earth Observation Data
- Text-Guided Semantic Image Encoder
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- PromptMoG: Enhancing Diversity in Long-Prompt Image Generation via Prompt Embedding Mixture-of-Gaussian Sampling
- Boosting Reasoning in Large Multimodal Models via Activation Replay
- MAPS: Preserving Vision-Language Representations via Module-Wise Proximity Scheduling for Better Vision-Language-Action Generalization
- Temporal-Visual Semantic Alignment: A Unified Architecture for Transferring Spatial Priors from Vision Models to Zero-Shot Temporal Tasks
- Unifying Perception and Action: A Hybrid-Modality Pipeline with Implicit Visual Chain-of-Thought for Robotic Action Generation
- SFA: Scan, Focus, and Amplify toward Guidance-aware Answering for Video TextVQA
- RADSeg: Unleashing Parameter and Compute Efficient Zero-Shot Open-Vocabulary Segmentation Using Agglomerative Models
- fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding
- On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
- UniGame: Turning a Unified Multimodal Model Into Its Own Adversary
- UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval
- ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- TRANSPORTER: Transferring Visual Semantics from VLM Manifolds
- Concept Regions Matter: Benchmarking CLIP with a New Cluster-Importance Approach
- Muskie: Multi-view Masked Image Modeling for 3D Vision Pre-training
- EchoVLA: Synergistic Declarative Memory for VLA-Driven Mobile Manipulation
- VK-Det: Visual Knowledge Guided Prototype Learning for Open-Vocabulary Aerial Object Detection
- CUS-GS: A Compact Unified Structured Gaussian Splatting Framework for Multimodal Scene Representation
- FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning
- HALO: High-Altitude Language-Conditioned Monocular Aerial Exploration and Navigation
- Downscaling Intelligence: Exploring Perception and Reasoning Bottlenecks in Small Multimodal Models
- Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- A Little More Like This: Text-to-Image Retrieval with Vision-Language Models Using Relevance Feedback
- Pillar-0: A New Frontier for Radiology Foundation Models
- METIS: Multi-Source Egocentric Training for Integrated Dexterous Vision-Language-Action Model
- SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
- Solving Spatial Supersensing Without Spatial Supersensing
- TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
- TOFA: Training-Free One-Shot Federated Adaptation for Vision-Language Models
- Upsample Anything: A Simple and Hard to Beat Baseline for Feature Upsampling
- BioBench: A Blueprint to Move Beyond ImageNet for Scientific ML Benchmarks
- Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval
- LLaVA3: Representing 3D Scenes like a Cubist Painter to Boost 3D Scene Understanding of VLMs
- X-WIN: Building Chest Radiograph World Model via Predictive Sensing
- EBind: a practical approach to space binding
- TTA: Transcribe, Translate and Alignment for Cross-lingual Speech Representation
- Find the Leak, Fix the Split: Cluster-Based Method to Prevent Leakage in Video-Derived Datasets
- Language-Guided Invariance Probing of Vision-Language Models
- Large Language Models Meet Extreme Multi-label Classification: Scaling and Multi-modal Framework
- Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization
- PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Analyzing Sustainability Messaging in Large-Scale Corporate Social Media
- Multimodal Large Language Models as Image Classifiers
- OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- SpaceVLM: Sub-Space Modeling of Negation in Vision-Language Models
- RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
- GCAgent: Long-Video Understanding via Schematic and Narrative Episodic Memory
- FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
- Rethinking Progression of Memory State in Robotic Manipulation: An Object-Centric Perspective
- Φeat: Physically Grounded Material Feature Representation
- Reinforcing Trustworthiness in Multimodal Emotional Support Systems
- Phantom Menace: Exploring and Enhancing the Robustness of VLA Models Against Physical Sensor Attacks
- SemanticVLA: Semantic-Aligned Sparsification and Enhancement for Efficient Robotic Manipulation
- Improving Perturbation-based Explanations by Understanding the Role of Uncertainty Calibration
- Audio-VLA: Adding Contact Audio Perception to Vision-Language-Action Model for Robotic Manipulation
- From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
- Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
- SliderEdit: Continuous Image Editing with Fine-Grained Instruction Control
- NeuCLIP: Efficient Large-Scale CLIP Training with Neural Normalizer Optimization
- Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking
- oboro: Text-to-Image Synthesis on Limited Data using Flow-based Diffusion Transformer with MMH Attention
- NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
- Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- ClusterMine: Robust Label-Free Visual Out-Of-Distribution Detection via Concept Mining from Text Corpora
- From Pretrain to Pain: Adversarial Vulnerability of Video Foundation Models Without Task Knowledge
- SlotVLA: Towards Modeling of Object-Relation Representations in Robotic Manipulation
- HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
- How Do VLAs Effectively Inherit from VLMs?
- 10 Open Challenges Steering the Future of Vision-Language-Action Models
- MIMIC-SR-ICD11: A Dataset for Narrative-Based Diagnosis
- EveryDayVLA: A Vision-Language-Action Model for Affordable Robotic Manipulation
- Another BRIXEL in the Wall: Towards Cheaper Dense Features
- Role-SynthCLIP: A Role Play Driven Diverse Synthetic Data Approach
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- Seeing Straight: Document Orientation Detection for Efficient OCR
- Seeing What You Say: Expressive Image Generation from Speech
- Fine-Tuning Vision-Language Models for Multimodal Polymer Property Prediction
- SCALE-VLP: Soft-Weighted Contrastive Volumetric Vision-Language Pre-training with Spatial-Knowledge Semantics
- Weakly Supervised Concept Learning with Class-Level Priors for Interpretable Medical Diagnosis
- VCode: a Multimodal Coding Benchmark with SVG as Symbolic Visual Representation
- XR-1: Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
- LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation
- V-Agent: An Interactive Video Search System Using Vision-Language Models
- Wave-Particle (Continuous-Discrete) Dualistic Visual Tokenization for Unified Understanding and Generation
- ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
- UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings
- Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
- Foundation Models for Trajectory Planning in Autonomous Driving: A Review of Progress and Open Challenges
- RzenEmbed: Towards Comprehensive Multimodal Retrieval
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- RegionRAG: Region-level Retrieval-Augmented Generation for Visual Document Understanding
- LongCat-Flash-Omni Technical Report
- Masked Diffusion Captioning for Visual Feature Learning
- Running VLAs at Real-time Speed
- Emu3.5: Native Multimodal Models are World Learners
- LoCoT2V-Bench: A Benchmark for Long-Form and Complex Text-to-Video Generation
- MV-MLM: Bridging Multi-View Mammography and Language for Breast Cancer Diagnosis and Risk Prediction
- Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
- Instance-Level Composed Image Retrieval
- DiagramEval: Evaluating LLM-Generated Diagrams via Graphs
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- Don't Blind Your VLA: Aligning Visual Representations for OOD Generalization
- Beyond Facial Consistency: Personalized Person Image Generation with Holistic Identity Preservation
- SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- MM-Nav: Multi-View VLA Model for Robust Visual Navigation via Multi-Expert Learning
- Can neurons speak? Semantic narration of vision at single-cell resolution
- SpectraDINO: Modality-Conditioned Adaptation of RGB Vision Foundation Models Across Infrared Bands
- Is Dimensionality a Barrier for Retrieval Models?
- SPROUT: A Scalable Diffusion Foundation Model for Agricultural Vision
- RL makes MLLMs see better than SFT
- EDVD-LLaMA: Explainable Deepfake Video Detection via Multimodal Large Language Model Reasoning
- Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug Design
- NanoVLA: Routing Decoupled Vision-Language Understanding for Nano-sized Generalist Robotic Policies
- BLM1: A Boundless Large Model for Cross-Space, Cross-Task, and Cross-Embodiment Learning
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- Improving Visual Discriminability of CLIP for Training-Free Open-Vocabulary Semantic Segmentation
- MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation
- RoboOmni: Proactive Robot Manipulation in Omni-modal Context
- A Survey on Efficient Vision-Language-Action Models
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- Solar flare forecasting with foundational transformer models across image, video, and time-series modalities
- DecoDINO: 3D Human-Scene Contact Prediction with Semantic Classification
- HyPerNav: Hybrid Perception for Object-Oriented Navigation in Unknown Environment
- WAON: Large-Scale and High-Quality Japanese Image-Text Pair Dataset for Vision-Language Models
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
- Automated Detection of Visual Attribute Reliance with a Self-Reflective Agent
- OpenHype: Hyperbolic Embeddings for Hierarchical Open-Vocabulary Radiance Fields
- SafetyPairs: Isolating Safety Critical Image Features with Counterfactual Image Generation
- Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
- FieldGen: From Teleoperated Pre-Manipulation Trajectories to Field-Guided Data Generation
- C-NAV: Towards Self-Evolving Continual Object Navigation in Open World
- GranViT: A Fine-Grained Vision Model With Autoregressive Perception For MLLMs
- Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
- COS3D: Collaborative Open-Vocabulary 3D Segmentation
- MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval
- BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
- Exploring Conditions for Diffusion models in Robotic Control
- Class-Aware Prototype Learning with Negative Contrast for Test-Time Adaptation of Vision-Language Models
- TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
- A Matter of Time: Revealing the Structure of Time in Vision-Language Models
- The Intricate Dance of Prompt Complexity, Quality, Diversity, and Consistency in T2I Models
- GigaBrain-0: A World Model-Powered Vision-Language-Action Model
- X-Ego: Acquiring Team-Level Tactical Situational Awareness via Cross-Egocentric Contrastive Video Representation Learning
- Semantic World Models
- Exploring a Unified Vision-Centric Contrastive Alternatives on Multi-Modal Web Documents
- CovMatch: Cross-Covariance Guided Multimodal Dataset Distillation with Trainable Text Encoder
- Automated urban waterlogging assessment and early warning through a mixture of foundation models
- ProLAP: Probabilistic Language-Audio Pre-Training
- OmniNWM: Omniscient Driving Navigation World Models
- StreamingTOM: Streaming Token Compression for Efficient Video Understanding
- Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- VFM-VAE: Vision Foundation Models Can Be Good Tokenizers for Latent Diffusion Models
- SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
- ReefNet: A Large scale, Taxonomically Enriched Dataset and Benchmark for Hard Coral Classification
- EgMM-Corpus: A Multimodal Vision-Language Dataset for Egyptian Culture
- Train a Unified Multimodal Data Quality Classifier with Synthetic Data
- Comprehensive language-image pre-training for 3D medical image understanding
- From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Learning an Image Editing Model without Image Editing Pairs
- WithAnyone: Towards Controllable and ID Consistent Image Generation
- DialectGen: Benchmarking and Improving Dialect Robustness in Multimodal Generation
- VLA2: Empowering Vision-Language-Action Models with an Agentic Framework for Unseen Concept Manipulation
- You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
- QDepth-VLA: Quantized Depth Prediction as Auxiliary Supervision for Vision-Language-Action Models
- Free-Grained Hierarchical Recognition
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- Exploratory Causal Inference in SAEnce
- When Embedding Models Meet: Procrustes Bounds and Applications
- Language as a Label: Zero-Shot Multimodal Classification of Everyday Postures under Data Scarcity
- Improving Visual Recommendation on E-commerce Platforms Using Vision-Language Models
- Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- Reasoning in Space via Grounding in the World
- UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
- VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- CompoDistill: Attention Distillation for Compositional Reasoning in Multimodal LLMs
- ImageSentinel: Protecting Visual Datasets from Unauthorized Retrieval-Augmented Image Generation
- SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model
- Diffusion Transformers with Representation Autoencoders
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Scaling Language-Centric Omnimodal Representation Learning
- How many samples to label for an application given a foundation model? Chest X-ray classification study
- Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
- Neural Weight Compression for Language Models
- Data or Language Supervision: What Makes CLIP Better than DINO?
- Embedding the Teacher: Distilling vLLM Preferences for Scalable Image Retrieval
- FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model
- Image-to-Video Transfer Learning based on Image-Language Foundation Models: A Comprehensive Survey
- Equipping Vision Foundation Model with Mixture of Experts for Out-of-Distribution Detection
- VITA-VLA: Efficiently Teaching Vision-Language Models to Act via Action Expert Distillation
- MRMR: A Realistic and Expert-Level Multidisciplinary Benchmark for Reasoning-Intensive Multimodal Retrieval
- Modeling Time-Lapse Trajectories to Characterize Cranberry Growth
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- NavSpace: How Navigation Agents Follow Spatial Intelligence Instructions
- MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- Enhancing Visual Prompting through Expanded Transformation Space and Overfitting Mitigation
- Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
- Approximate Domain Unlearning for Vision-Language Models
- NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
- Don't Run with Scissors: Pruning Breaks VLA Models but They Can Be Recovered
- On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- In-Context Clustering with Large Language Models
- TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
- TrackVLA++: Unleashing Reasoning and Memory Capabilities in VLA Models for Embodied Visual Tracking
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- Revisiting Mixout: An Overlooked Path to Robust Finetuning
- Growing Visual Generative Capacity for Pre-Trained MLLMs
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- TTRV: Test-Time Reinforcement Learning for Vision Language Models
- Heptapod: Language Modeling on Visual Signals
- Ming-UniVision: Joint Image Understanding and Generation with a Unified Continuous Tokenizer
- From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
- VUGEN: Visual Understanding priors for GENeration
- SIGMA-GEN: Structure and Identity Guided Multi-subject Assembly for Image Generation
- Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
- VCoT-Grasp: Grasp Foundation Models with Visual Chain-of-Thought Reasoning for Language-driven Grasp Generation
- Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- Do You Know Where Your Camera Is? View-Invariant Policy Learning with Camera Conditioning
- HyperVLA: Efficient Inference in Vision-Language-Action Models via Hypernetworks
- Visual Representations inside the Language Model
- Activation Quantization of Vision Encoders Needs Prefixing Registers
- Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Personalizing Retrieval using Joint Embeddings or "the Return of Fluffy"
- Think Then Embed: Generative Context Improves Multimodal Embedding
- Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
- ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- FrameOracle: Learning What to See and How Much to See in Videos
- From Scope to Script: An Automated Report Generation Model for Gastrointestinal Endoscopy
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- Improving GUI Grounding with Explicit Position-to-Coordinate Mapping
- Confidence and Dispersity as Signals: Unsupervised Model Evaluation and Ranking
- SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval
- Visual Language Model as a Judge for Object Detection in Industrial Diagrams
- ModernVBERT: Towards Smaller Visual Document Retrievers
- Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
- Hybrid Training for Vision-Language-Action Models
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- Can World Models Benefit VLMs for World Dynamics?
- Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
- Optimizing What Matters: AUC-Driven Learning for Robust Neural Retrieval
- MLA: A Multisensory Language-Action Model for Multimodal Understanding and Forecasting in Robotic Manipulation
- TimeScope: Towards Task-Oriented Temporal Grounding In Long Videos
- FinGAN: An Interpretable RSS Generation Network for Scalable Fingerprint Localization
- ProfVLM: A Lightweight Video-Language Model for Multi-View Proficiency Estimation
- DGM4+: Dataset Extension for Global Scene Inconsistency
- SGS: Segmentation-Guided Scoring for Global Scene Inconsistencies
- Learning Egocentric In-Hand Object Segmentation through Weak Supervision from Human Narrations
- The Impact of Scaling Training Data on Adversarial Robustness
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- From Perception to Cognition: A Survey of Vision-Language Interactive Reasoning in Multimodal Large Language Models
- Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
- Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- VT-FSL: Bridging Vision and Text with LLMs for Few-Shot Learning
- StreamForest: Efficient Online Video Understanding with Persistent Event Memory
- AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy
- When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- NeMo: Needle in a Montage for Video-Language Understanding
- FreeRet: MLLMs as Training-Free Retrievers
- A TRIANGLE Enables Multimodal Alignment Beyond Cosine Similarity
- EduVidQA: Generating and Evaluating Long-form Answers to Student Questions based on Lecture Videos
- CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
- Uni4D-LLM: A Unified SpatioTemporal-Aware VLM for 4D Understanding and Generation
- UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
- Focusing on What Matters: Object-Agent-centric Tokenization for Vision Language Action models
- Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- RAVEN: Resilient Aerial Navigation via Open-Set Semantic Memory and Behavior Adaptation
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Detecting YouTube Scam Videos via Multimodal Signals and Policy Reasoning
- DentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice
- Preventing Robotic Jailbreaking via Multimodal Domain Adaptation
- Transferring Vision-Language-Action Models to Industry Applications: Architectures, Performance, and Challenges
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- Reinforcement Learning-Based Prompt Template Stealing for Text-to-Image Models
- Multilingual Vision-Language Models, A Survey
- Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data
- Tiny but Mighty: A Software-Hardware Co-Design Approach for Efficient Multimodal Inference on Battery-Powered Small Devices
- Un-Doubling Diffusion: LLM-guided Disambiguation of Homonym Duplication
- Less Precise Can Be More Reliable: A Systematic Evaluation of Quantization's Impact on VLMs Beyond Accuracy
- SupCLAP: Controlling Optimization Trajectory Drift in Audio-Text Contrastive Learning with Support Vector Regularization
- WAVECLIP: Wavelet Tokenization for Adaptive-Resolution CLIP
- Concepts in Motion: Temporal Concept Bottleneck Model for Interpretable Video Classification
- Generalist Robot Manipulation beyond Action Labeled Data
- CHURRO: Making History Readable with an Open-Weight Large Vision-Language Model for High-Accuracy, Low-Cost Historical Text Recognition
- Personality Vector: Modulating Personality of Large Language Models by Model Merging
- Interpreting ResNet-based CLIP via Neuron-Attention Decomposition
- LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA
- Foundation Models for Face Presentation Attack Detection: A Unified Linear-Probing Benchmark
- MeshFM: 2D Features Are All You Need for 3D Shape Understanding
- SemAnCorr: Semantic Anchored Correspondence for Zero-Shot Manipulation Skill Transfer
- VisualRouter: Query-Grounded Visual Sampling for Long Video Understanding
- Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer
- AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
- Learning from Compressed CT: Feature Attention Style Transfer and Structured Factorized Projections for Resource-Efficient Medical Image Analysis
- MolmoB0T: Large-Scale Simulation Enables Zero-Shot Manipulation
- Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
- AVAM: Universal Training-free Adaptive Visual Anchoring Embedded into Multimodal Large Language Model for Multi-image Question Answering
- Long Story Short: Disentangling Compositionality and Long-Caption Understanding in VLMs
- Bi-VLA: Bilateral Control-Based Imitation Learning via Vision-Language Fusion for Action Generation
- Benchmarking Vision-Language and Multimodal Large Language Models in Zero-shot and Few-shot Scenarios: A study on Christian Iconography
- Self-Alignment Learning to Improve Myocardial Infarction Detection from Single-Lead ECG
- Global Minimizers of Sigmoid Contrastive Loss
- VGGT-DP: Generalizable Robot Control via Vision Foundation Models
- Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
- Latent Action Pretraining Through World Modeling
- Visual Instruction Pretraining for Domain-Specific Foundation Models
- Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
- ComposeMe: Attribute-Specific Image Prompts for Controllable Human Image Generation
- MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
- Learning Contrastive Multimodal Fusion with Improved Modality Dropout for Disease Detection and Prediction
- I-FailSense: Towards General Robotic Failure Detection with Vision-Language Models
- Efficient Long-Tail Learning in Latent Space by sampling Synthetic Data
- The SAGES Critical View of Safety Challenge: A Global Benchmark for AI-Assisted Surgical Quality Assessment
- KV-Efficient VLA: A Method to Speed up Vision Language Models with RNN-Gated Chunked KV Cache
- Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
- ADVEDM:Fine-grained Adversarial Attack against VLM-based Embodied Agents
- AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- GP3: A 3D Geometry-Aware Policy with Multi-View Images for Robotic Manipulation
- UNIV: Unified Foundation Model for Infrared and Visible Modalities
- PCSR: Pseudo-label Consistency-Guided Sample Refinement for Noisy Correspondence Learning
- FiLM-Nav: Efficient and Generalizable Navigation via VLM Fine-tuning
- Dynamic Classifier-Free Diffusion Guidance via Online Feedback
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Multimodal Representation Learning Conditioned on Semantic Relations
- PRISM: Product Retrieval In Shopping Carts using Hybrid Matching
- CollabVLA: Self-Reflective Vision-Language-Action Model Dreaming Together with Human
- Frame Sampling Strategies Matter: A Benchmark for small vision language models
- Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
- Toward Embodiment Equivariant Vision-Language-Action Policy
- AToken: A Unified Tokenizer for Vision
- Constrained Prompt Enhancement for Improving Zero-Shot Generalization of Vision-Language Models
- Dense Video Understanding with Gated Residual Tokenization
- CLAW: A Vision-Language-Action Framework for Weight-Aware Robotic Grasping
- SeqVLA: Sequential Task Execution for Long-Horizon Manipulation with Completion-Aware Vision-Language-Action Model
- GeoAware-VLA: Implicit Geometry Aware Vision-Language-Action Model
- An Empirical Analysis of VLM-based OOD Detection: Mechanisms, Advantages, and Sensitivity
- From Embeddings to Equations: Genetic-Programming Surrogates for Interpretable Transformer Classification
- Brought a Gun to a Knife Fight: Modern VFM Baselines Outgun Specialized Detectors on In-the-Wild AI Image Detection
- What Makes a Good Generated Image? Investigating Human and Multimodal LLM Image Preference Alignment
- Hunyuan3D Studio: End-to-End AI Pipeline for Game-Ready 3D Asset Generation
- The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
- Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
- Embodied Navigation Foundation Model
- Zero-shot Multimodal Document Retrieval via Cross-modal Question Generation
- How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
- Synthetic vs. Real Training Data for Visual Navigation
- Pathological Truth Bias in Vision-Language Models
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- F4-ITS: Fine-grained Feature Fusion for Food Image-Text Search
- InPhyRe Discovers: Large Multimodal Models Struggle in Inductive Physical Reasoning
- Towards Understanding Visual Grounding in Visual Language Models
- Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution
- VLA-Adapter: An Effective Paradigm for Tiny-Scale Vision-Language-Action Model
- Image Recognition with Vision and Language Embeddings of VLMs
- Chirality in Action: Time-Aware Video Representation Learning by Latent Straightening
- Recurrence Meets Transformers for Universal Multimodal Retrieval
- SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
- Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
- Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data
- Index-Preserving Lightweight Token Pruning for Efficient Document Understanding in Vision-Language Models
- Text4Seg++: Advancing Image Segmentation via Generative Language Modeling
- F1: A Vision-Language-Action Model Bridging Understanding and Generation to Actions
- SVGauge: Towards Human-Aligned Evaluation for SVG Generation
- MV-RAG: Retrieval Augmented Multiview Diffusion
- SpecPrune-VLA: Accelerating Vision-Language-Action Models via Action-Aware Self-Speculative Pruning
- SGS-3D: High-Fidelity 3D Instance Segmentation via Reliable Semantic Mask Splitting and Growing
- Symbolic Graphics Programming with Large Language Models
- Guideline-Consistent Segmentation via Multi-Agent Refinement
- Not All Splits Are Equal: Rethinking Attribute Generalization Across Unrelated Categories
- SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
- Weakly-Supervised Learning of Dense Functional Correspondences
- Modular Embedding Recomposition for Incremental Learning
- Skywork UniPic 2.0: Building Kontext Model with Online RL for Unified Multimodal Model
- E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition
- OneCAT: Decoder-Only Auto-Regressive Model for Unified Understanding and Generation
- Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
- CMRAG: Co-modality-based visual document retrieval and question answering
- Fidelity-preserving enhancement of ptychography with foundational text-to-image models
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Measuring Image-Relation Alignment: Reference-Free Evaluation of VLMs and Synthetic Pre-training for Open-Vocabulary Scene Graph Generation
- Do Video Language Models Really Know Where to Look? Diagnosing Attention Failures in Video Language Models
- Street-Level Geolocalization Using Multimodal Large Language Models and Retrieval-Augmented Generation
- OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
- Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- Beyond Pixels: Introducing Geometric-Semantic World Priors for Video-based Embodied Models via Spatio-temporal Alignment
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- Boosting Pathology Foundation Models via Few-shot Prompt-tuning for Rare Cancer Subtyping
- BiListing: Modality Alignment for Listings
- StreamMem: Query-Agnostic KV Cache Memory for Streaming Video Understanding
- MobileCLIP2: Improving Multi-Modal Reinforced Training
- Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies
- SCAR: A Characterization Scheme for Multi-Modal Dataset
- MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- DeepMEL: A Multi-Agent Collaboration Framework for Multimodal Entity Linking
- USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning
- Survey of Vision-Language-Action Models for Embodied Manipulation
- On Evaluating the Adversarial Robustness of Foundation Models for Multimodal Entity Linking
- CLIPSym: Delving into Symmetry Detection with CLIP
- RotBench: Evaluating Multimodal Large Language Models on Identifying Image Rotation
- Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
- CAST: Counterfactual Labels Improve Instruction Following in Vision-Language-Action Models
- Mitigating Easy Option Bias in Multiple-Choice Question Answering
- Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
- Learning to Steer: Input-dependent Steering for Multimodal LLMs
- Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
- Cross-Domain Few-Shot Learning via Multi-View Collaborative Optimization with Vision-Language Models
- DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
- Infusing fine-grained visual knowledge to Vision-Language Models
- UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- EVTP-IVS: Effective Visual Token Pruning For Unifying Instruction Visual Segmentation In Multi-Modal Large Language Models
- OVSegDT: Segmenting Transformer for Open-Vocabulary Object Goal Navigation
- Generalized Decoupled Learning for Enhancing Open-Vocabulary Dense Perception
- Processing and acquisition traces in visual encoders: What does CLIP know about your camera?
- CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
- ReconVLA: Reconstructive Vision-Language-Action Model as Effective Robot Perceiver
- Failures to Surface Harmful Contents in Video Large Language Models
- Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation
- LLMC+: Benchmarking Vision-Language Model Compression with a Plug-and-play Toolkit
- On the dynamic evolution of CLIP texture-shape bias and its relationship to human alignment and model robustness
- MoIIE: Mixture of Intra- and Inter-Modality Experts for Large Vision Language Models
- SHREC 2025: Retrieval of Optimal Objects for Multi-modal Enhanced Language and Spatial Assistance (ROOMELSA)
- OmniVTLA: Vision-Tactile-Language-Action Model with Semantic-Aligned Tactile Sensing
- Closing the Performance Gap in Generative Recommenders with Collaborative Tokenization and Efficient Modeling
- Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
- ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
- Being-M0.5: A Real-Time Controllable Vision-Language-Motion Model
- VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
- LET-US: Long Event-Text Understanding of Scenes
- AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
- BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- BiXSE: Improving Dense Retrieval via Probabilistic Graded Relevance Distillation
- Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
- CountQA: How Well Do MLLMs Count in the Wild?
- Bifrost-1: Bridging Multimodal LLMs and Diffusion Models with Patch-level CLIP Latents
- Q-CLIP: Unleashing the Power of Vision-Language Models for Video Quality Assessment through Unified Cross-Modal Adaptation
- LATTE: Learning Aligned Transactions and Textual Embeddings for Bank Clients
- Adapting Vision-Language Models Without Labels: A Comprehensive Survey
- Explaining Similarity in Vision-Language Encoders with Weighted Banzhaf Interactions
- VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
- Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
- X-SAM: From Segment Anything to Any Segmentation
- FrEVL: Leveraging Frozen Pretrained Embeddings for Efficient Vision-Language Understanding
- Composed Object Retrieval: Object-level Retrieval via Composed Expressions
- Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting
- What Holds Back Open-Vocabulary Segmentation?
- OpenLifelogQA: An Open-Ended Multi-Modal Lifelog Question-Answering Dataset
- VLMQ: Efficient Post-Training Quantization for Large Vision-Language Models via Hessian Augmentation
- Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation
- UniEdit-I: Training-free Image Editing for Unified VLM via Iterative Understanding, Editing and Verifying
- XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs
- VITRIX-CLIPIN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions
- VeOmni: Scaling Any Modality Model Training with Model-Centric Distributed Recipe Zoo
- Patho-AgenticRAG: Towards Multimodal Agentic Retrieval-Augmented Generation for Pathology VLMs via Reinforcement Learning
- S-RRG-Bench: Structured Radiology Report Generation with Fine-Grained Evaluation Framework
- RICL: Adding In-Context Adaptability to Pre-Trained Vision-Language-Action Models
- Multimodal Large Language Models for End-to-End Affective Computing: Benchmarking and Boosting with Generative Knowledge Prompting
- TimeExpert: An Expert-Guided Video LLM for Video Temporal Grounding
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
- A Large-Scale Benchmark of Cross-Modal Learning for Histology and Gene Expression in Spatial Transcriptomics
- SaviorRec: Semantic-Behavior Alignment for Cold-Start Recommendation
- YOLO-Count: Differentiable Object Counting for Text-to-Image Generation
- HiPrune: Training-Free Visual Token Pruning via Hierarchical Attention in Vision-Language Models
- From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model
- SUB: Benchmarking CBM Generalization via Synthetic Attribute Substitutions
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- H-RDT: Human Manipulation Enhanced Bimanual Robotic Manipulation
- MoCHA: Advanced Vision-Language Reasoning with MoE Connector and Hierarchical Group Attention
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
- SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
Discussions
Related