Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
2024/06/24 by Shengbang Tong, Ellis Brown, Tong, Shengbang +26 · 2 voices · 322 citations
Computer Science · #Advanced Computational Techniques and Applications #cs.CV
paper · pdf · doi:10.48550/arxiv.2406.16860
arxiv published 2024/06/24 · arxiv updated 2024/12/04
Abstract
We introduce Cambrian-1, a family of multimodal LLMs (MLLMs) designed with a vision-centric approach. While stronger language models can enhance multimodal capabilities, the design choices for vision components are often insufficiently explored and disconnected from visual representation learning research. This gap hinders accurate sensory grounding in real-world scenarios. Our study uses LLMs and visual instruction tuning as an interface to evaluate various visual representations, offering new insights into different models and architectures -- self-supervised, strongly supervised, or combinations thereof -- based on experiments with over 20 vision encoders. We critically examine existing MLLM benchmarks, address the difficulties involved in consolidating and interpreting results from various tasks, and introduce a new vision-centric benchmark, CV-Bench. To further improve visual grounding, we propose the Spatial Vision Aggregator (SVA), a dynamic and spatially-aware connector that integrates high-resolution vision features with LLMs while reducing the number of tokens. Additionally, we discuss the curation of high-quality visual instruction-tuning data from publicly available sources, emphasizing the importance of data source balancing and distribution ratio. Collectively, Cambrian-1 not only achieves state-of-the-art performance but also serves as a comprehensive, open cookbook for instruction-tuned MLLMs. We provide model weights, code, supporting tools, datasets, and detailed instruction-tuning and evaluation recipes. We hope our release will inspire and accelerate advancements in multimodal systems and visual representation learning.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- The Spatial Blindspot of Vision-Language Models
- Self-Boosting Vision-Language Models with Noisy Student On-Policy Self-Distillation
- Unified Embodied VLM Reasoning with Robotic Action via Autoregressive Discretized Pre-training
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
- Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
- A Benchmark for Ultra-High-Resolution Remote Sensing MLLMs
- Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
- N3D-VLM: Native 3D Grounding Enables Accurate Spatial Reasoning in Vision-Language Models
- Seeing Beyond Words: Self-Supervised Visual Learning for Multimodal Large Language Models
- PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
- V-REX: Benchmarking Exploratory Visual Reasoning via Chain-of-Questions
- JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction
- Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
- DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
- Exploring MLLM-Diffusion Information Transfer with MetaCanvas
- SmokeBench: Evaluating Multimodal Large Language Models for Wildfire Smoke Detection
- Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
- Mull-Tokens: Modality-Agnostic Latent Thinking
- No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- Stitch and Tell: A Structured Multimodal Data Augmentation Method for Spatial Understanding
- Unleashing the Intrinsic Visual Representation Capability of Multimodal Large Language Models
- Knowing the Answer Isn't Enough: Fixing Reasoning Path Failures in LVLMs
- Towards Cross-View Point Correspondence in Vision-Language Models
- Jina-VLM: Small Multilingual Vision Language Model
- Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models
- PAI-Bench: A Comprehensive Benchmark For Physical AI
- DialBench: Towards Accurate Reading Recognition of Pointer Meter using Large Foundation Models
- From Pixels to Feelings: Aligning MLLMs with Human Cognitive Perception of Images
- Geometrically-Constrained Agent for Spatial Reasoning
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition
- EM-KD: Distilling Efficient Multimodal Large Language Model with Unbalanced Vision Tokens
- Text-Guided Semantic Image Encoder
- Vision-Language Memory for Spatial Reasoning
- Thinking in 360°: Humanoid Visual Search in the Wild
- VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering
- LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
- Perceptual Taxonomy: Evaluating and Guiding Hierarchical Scene Reasoning in Vision-Language Models
- Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens
- SO-Bench: A Structural Output Evaluation of Multimodal LLMs
- When Better Teachers Don't Make Better Students: Revisiting Knowledge Distillation for CLIP Models in VQA
- Attention Guided Alignment in Efficient Vision-Language Models
- IndustryNav: Exploring Spatial Reasoning of Embodied Agents in Dynamic Industrial Navigation
- ChainV: Atomic Visual Hints Make Multimodal Reasoning Shorter and Better
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Is your VLM Sky-Ready? A Comprehensive Spatial Intelligence Benchmark for UAV Navigation
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
- Direct Visual Grounding by Directing Attention of Visual Tokens
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- Multimodal Large Language Models for Low-Resource Languages: A Case Study for Basque
- Simple Vision-Language Math Reasoning via Rendered Text
- Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- Revisiting the Data Sampling in Multimodal Post-training from a Difficulty-Distinguish View
- iFlyBot-VLM Technical Report
- Visual Spatial Tuning
- Cambrian-S: Towards Spatial Supersensing in Video
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Benchmark Designers Should "Train on the Test Set" to Expose Exploitable Non-Visual Shortcuts
- IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- Seeing Straight: Document Orientation Detection for Efficient OCR
- TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- Masked Diffusion Captioning for Visual Feature Learning
- Do Vision-Language Models Measure Up? Benchmarking Visual Measurement Reading with MeasureBench
- Multimodal Spatial Reasoning in the Large Model Era: A Survey and Benchmarks
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language Models
- InstructPLM-mu: 1-Hour Fine-Tuning of ESM2 Beats ESM3 in Protein Mutation Predictions
- Mitigating Modal Imbalance in Multimodal Reasoning
- VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
- RL makes MLLMs see better than SFT
- SafeVision: Efficient Image Guardrail with Robust Policy Adherence and Explainability
- SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning
- UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
- Dexbotic: Open-Source Vision-Language-Action Toolbox
- MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- EmbodiedBrain: Expanding Performance Boundaries of Task Planning for Embodied Intelligence
- Data-Centric Lessons To Improve Speech-Language Pretraining
- Grasp Any Region: Towards Precise, Contextual Pixel Understanding for Multimodal LLMs
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views
- FineVision: Open Data Is All You Need
- MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
- Vision-Centric Activation and Coordination for Multimodal Large Language Models
- Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
- Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Scope: Selective Cross-modal Orchestration of Visual Perception Experts
- Point Prompting: Counterfactual Tracking with Video Diffusion Models
- Scaling Language-Centric Omnimodal Representation Learning
- A Survey on Agentic Multimodal Large Language Models
- Data or Language Supervision: What Makes CLIP Better than DINO?
- RefineShot: Rethinking Cinematography Understanding with Foundational Skill Evaluation
- Task-Aware Resolution Optimization for Visual Large Language Models
- SpaceVista: All-Scale Visual Spatial Reasoning from mm to km
- Unleashing Perception-Time Scaling to Multimodal Reasoning Models
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation
- SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Evaluating Small Vision-Language Models on Distance-Dependent Traffic Perception
- TIGeR: Tool-Integrated Geometric Reasoning in Vision-Language Models for Robotics
- Visual Representations inside the Language Model
- MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
- OneFlow: Concurrent Mixed-Modal and Interleaved Generation with Edit Flows
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
- Mitigating Hallucination in Multimodal LLMs with Layer Contrastive Decoding
- Vision Function Layer in Multimodal LLMs
- Multi-modal Data Spectrum: Multi-modal Datasets are Multi-dimensional
- Understanding Language Prior of LVLMs by Contrasting Chain-of-Embedding
- MultiMat: Multimodal Program Synthesis for Procedural Materials using Large Multimodal Models
- The Photographer Eye: Teaching Multimodal Large Language Models to Understand Image Aesthetics like Photographers
- Catching the Details: Self-Distilled RoI Predictors for Fine-Grained MLLM Perception
- History-Aware Visuomotor Policy Learning via Point Tracking
- MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
- Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
- SAIL-VL2 Technical Report
- Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
- ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
- Measuring Epistemic Humility in Multimodal Large Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- RoboChemist: Long-Horizon and Safety-Compliant Robotic Chemical Experimentation
- Visual Representation Alignment for Multimodal Large Language Models
- Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
- Category-level Text-to-Image Retrieval Improved: Bridging the Domain Gap with Diffusion Models and Vision Encoders
- MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
- GLSim: Detecting Object Hallucinations in LVLMs via Global-Local Similarity
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Unveiling Trust in Multimodal Large Language Models: Evaluation, Analysis, and Mitigation
- XFinBench: Benchmarking LLMs in Complex Financial Problem Solving and Reasoning
- JRDB-Reasoning: A Difficulty-Graded Benchmark for Visual Reasoning in Robotics
- LET-US: Long Event-Text Understanding of Scenes
- Training-Free Multimodal Large Language Model Orchestration
- MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning
- Large Language Models Facilitate Vision Reflection in Image Classification
- VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- TARS: MinMax Token-Adaptive Preference Strategy for Hallucination Reduction in MLLMs
- METEOR: Multi-Encoder Collaborative Token Pruning for Efficient Vision Language Models
- Region-based Cluster Discrimination for Visual Representation Learning
- OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
- From Intention to Execution: Probing the Generalization Boundaries of Vision-Language-Action Models
- AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation
- EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
- Better Reasoning with Less Data: Enhancing VLMs Through Unified Modality Scoring
- AdaBrain-Bench: Benchmarking Brain Foundation Models for Brain-Computer Interface Applications
- AVA-Bench: Atomic Visual Ability Benchmark for Vision Foundation Models
- Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
- M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models
- Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
- Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor
- Hierarchical Feature Alignment for Gloss-Free Sign Language Translation
- Omni-Video: Democratizing Unified Video Understanding and Generation
- LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Vision-Language Models Can't See the Obvious
- MODA: MOdular Duplex Attention for Multimodal Perception, Cognition, and Emotion Understanding
- Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
- Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
- How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
- GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- Empowering Small VLMs to Think with Dynamic Memorization and Exploration
- Task-Aware KV Compression For Cost-Effective Long Video Understanding
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology
- MedErr-CT: A Visual Question Answering Benchmark for Identifying and Correcting Errors in CT Reports
- HAWAII: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language Models
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs
- SpaCE-10: A Comprehensive Benchmark for Multimodal Large Language Models in Compositional Spatial Intelligence
- Dual-Stage Value-Guided Inference with Margin-Based Reward Adjustment for Fast and Faithful VLM Captioning
- Aligning Text, Images, and 3D Structure Token-by-Token
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- SpatialLM: Training Large Language Models for Structured Indoor Modeling
- Play to Generalize: Learning to Reason Through Game Play
- SAP-Bench: Benchmarking Multimodal Large Language Models in Surgical Action Planning
- Native Visual Understanding: Resolving Resolution Dilemmas in Vision-Language Models
- Domain Specific Benchmarks for Evaluating Multimodal Large Language Models
- Image Corruption-Inspired Membership Inference Attacks against Large Vision-Language Models
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation
- PAL: Probing Audio Encoders via LLMs -- Audio Information Transfer into LLMs
- RARL: Improving Medical VLM Reasoning and Generalization with Reinforcement Learning and LoRA under Data and Hardware Constraints
- Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models
- Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
- Refer to Any Segmentation Mask Group With Vision-Language Prompts
- MLLM-CL: Continual Learning for Multimodal Large Language Models
- X-Transfer Attacks: Towards Super Transferable Adversarial Attacks on CLIP
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
- DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding
- MiMo-VL Technical Report
- Vision Remember: Recovering Visual Information in Efficient LVLM with Vision Feature Resampling
- PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
- Seeing the Arrow of Time in Large Multimodal Models
- RoboEgo System Card: An Omnimodal Model with Native Full Duplexity
- MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
- Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement
- SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
- Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
- un2CLIP: Improving CLIP's Visual Detail Capturing Ability via Inverting unCLIP
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
- Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
- ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- OmniEarth-Bench: Towards Holistic Evaluation of Earth's Six Spheres and Cross-Spheres Interactions with Multimodal Observational Earth Data
- QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
- MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Jigsaw-R1: A Study of Rule-based Visual Reinforcement Learning with Jigsaw Puzzles
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- Spoken question answering for visual queries
- Zero-Shot Vision Encoder Grafting via LLM Surrogates
- Do You See Me : A Multidimensional Benchmark for Evaluating Visual Perception in Multimodal LLMs
- NegVQA: Can Vision Language Models Understand Negation?
- GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
- Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models
- FlowCut: Rethinking Redundancy via Information Flow for Efficient Vision-Language Models
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities
- Deformable Attentive Visual Enhancement for Referring Segmentation Using Vision-Language Model
- Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning
- Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs
- Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
- The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
- Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
- SophiaVL-R1: Reinforcing MLLMs Reasoning with Thinking Reward
- Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- R1-ShareVL: Incentivizing Reasoning Capability of Multimodal Large Language Models via Share-GRPO
- SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence
- Panoptic Captioning: An Equivalence Bridge for Image and Text
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
- Streamline Without Sacrifice -- Squeeze out Computation Redundancy in LMM
- Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning
- OViP: Online Vision-Language Preference Learning for VLM Hallucination
- Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs
- Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Modality-Balancing Preference Optimization of Large Multimodal Models by Adversarial Negative Mining
- Toward Embodied AGI: A Review of Embodied AI and the Road Ahead
- Visual Instruction Bottleneck Tuning
- Visionary-R1: Mitigating Shortcuts in Visual Reasoning with Reinforcement Learning
- Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
- ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models
- LogicOCR: Do Your Large Multimodal Models Excel at Logical Reasoning on Text-Rich Images?
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
- Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans
- Relax Within, Balance Across: Geometry-Guided Load Balancing for Vision-Language Mixture-of-Experts
- Privacy of Groups in Dense Street Imagery
- SITE: towards Spatial Intelligence Thorough Evaluation
- QTrack: Query-Driven Reasoning for Multi-modal MOT
- MiniMax Sparse Attention
- From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
- SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models
- COMPACT: COMPositional Atomic-to-Complex Visual Capability Tuning
- Beyond Language Modeling: An Exploration of Multimodal Pretraining
- Black-Box Visual Prompt Engineering for Mitigating Object Hallucination in Large Vision Language Models
- Multimodal Language Models See Better When They Look Shallower
- Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
- Let ViT Speak: Generative Language-Image Pre-training
- Revisiting Data Auditing in Large Vision-Language Models
- VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes
- RUTA: Principled Visual Token Allocation via Rate-Utility Optimization
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
- Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
- Scaling and Beyond: Advancing Spatial Reasoning in MLLMs Requires New Recipes
- LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
- AnomalyR1: A GRPO-based End-to-end MLLM for Industrial Anomaly Detection
- GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?
- Disentangling 3D Modeling from Spatial Reasoning
- PRISM: Priority-aware Rubric Internalization via Structured Multimodal Data Synthesis
- MIEB: Massive Image Embedding Benchmark
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single Transformer
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Multimodal Long Video Modeling Based on Temporal Dynamic Context
- AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
- Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
- TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models
- FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- How Can Objects Help Video-Language Understanding?
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
- ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- Kimi-VL Technical Report
- Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
- OmniCaptioner: One Captioner to Rule Them All
- Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
- OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
- SmolVLM: Redefining small and efficient multimodal models
- LEO-MINI: An Efficient Multimodal Large Language Model using Conditional Token Reduction and Mixture of Multi-Modal Experts
- TokenFLEX: Unified VLM Training for Flexible Visual Tokens Inference
Discussions
Related