Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
2024/09/25 by Matt Deitke, Deitke, Matt, Christopher Clark +100 · 4 voices · 287 citations
Computer Science · #Semantic Web and Ontologies #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2409.17146
Abstract
Today's most advanced vision-language models (VLMs) remain proprietary. The strongest open-weight models rely heavily on synthetic data from proprietary VLMs to achieve good performance, effectively distilling these closed VLMs into open ones. As a result, the community has been missing foundational knowledge about how to build performant VLMs from scratch. We present Molmo, a new family of VLMs that are state-of-the-art in their class of openness. Our key contribution is a collection of new datasets called PixMo, including a dataset of highly detailed image captions for pre-training, a free-form image Q&A dataset for fine-tuning, and an innovative 2D pointing dataset, all collected without the use of external VLMs. The success of our approach relies on careful modeling choices, a well-tuned training pipeline, and, most critically, the quality of our newly collected datasets. Our best-in-class 72B model not only outperforms others in the class of open weight and data models, but also outperforms larger proprietary models including Claude 3.5 Sonnet, and Gemini 1.5 Pro and Flash, second only to GPT-4o based on both academic benchmarks and on a large human evaluation. Our model weights, new datasets, and source code are available at https://molmo.allenai.org/blog.
Cited by
- Same or Not? Enhancing Visual Perception in Vision-Language Models
- CountGD++: Generalized Prompting for Open-World Counting
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone
- Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
- UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models
- The Spatial Blindspot of Vision-Language Models
- Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model
- CLIP Is Shortsighted: Paying Attention Beyond the First Sentence
- Latent Implicit Visual Reasoning
- Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- CodeDance: A Dynamic Tool-integrated MLLM for Executable Visual Reasoning
- Evaluating Large Language Models on Multimodal Chemistry Olympiad Exams
- SAGE: Training Smart Any-Horizon Agents for Long Video Reasoning with Reinforcement Learning
- Adapting MLLMs for Nuanced Video Retrieval
- β-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
- Grounding Everything in Tokens for Multimodal Large Language Models
- InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models
- Mind to Hand: Purposeful Robotic Control via Embodied Reasoning
- CoT4Det: A Chain-of-Thought Framework for Perception-Oriented Vision-Language Tasks
- RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension
- M4-RAG: A Massive-Scale Multilingual Multi-Cultural Multimodal RAG
- Towards Cross-View Point Correspondence in Vision-Language Models
- SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL
- Jina-VLM: Small Multilingual Vision Language Model
- DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling
- FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
- Artemis: Structured Visual Reasoning for Perception Policy Learning
- SocialFusion: Addressing Social Degradation in Pre-trained Vision-Language Models
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- Action-guided generation of 3D functionality segmentation data
- Obstruction reasoning for robotic grasping
- World in a Frame: Understanding Culture Mixing as a New Challenge for Vision-Language Models
- INSIGHT: An Interpretable Neural Vision-Language Framework for Reasoning of Generative Artifacts
- Multi-Crit: Benchmarking Multimodal Judges on Pluralistic Criteria-Following
- Qwen3-VL Technical Report
- LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Benchmarking Corruption Robustness of LVLMs: A Discriminative Benchmark and Robustness Alignment Metric
- FastMMoE: Accelerating Multimodal Large Language Models through Dynamic Expert Activation and Routing-Aware Token Pruning
- Don't Learn, Ground: A Case for Natural Language Inference with Visual Grounding
- BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Direct Visual Grounding by Directing Attention of Visual Tokens
- Uni-MoE-2.0-Omni: Scaling Language-Centric Omnimodal Large Model with Advanced MoE, Training and Data
- RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
- AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
- "It's trained by non-disabled people": Evaluating How Image Quality Affects Product Captioning with Vision-Language Models
- Multimodal LLMs Do Not Compose Skills Optimally Across Modalities
- SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards
- VLM-driven Skill Selection for Robotic Assembly Tasks
- iFlyBot-VLM Technical Report
- Long Grounded Thoughts: Synthesizing Visual Problems and Reasoning Chains at Scale
- DeepEyesV2: Toward Agentic Multimodal Model
- Visual Spatial Tuning
- SIMS-V: Simulated Instruction-Tuning for Spatial Video Understanding
- Multi-Task Learning for Visually Grounded Reasoning in Gastrointestinal VQA
- NVIDIA Nemotron Nano V2 VL
- Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
- LGCA: Enhancing Semantic Representation via Progressive Expansion
- LongCat-Flash-Omni Technical Report
- NaviTrace: Evaluating Embodied Navigation of Vision-Language Models
- Knowledge-Guided Textual Reasoning for Explainable Video Anomaly Detection via LLMs
- ChartNet: A Million-Scale, High-Quality Multimodal Dataset for Robust Chart Understanding
- TiPToP: A Modular Open-Vocabulary Robot Manipulation System That Plans
- PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement
- Conflict Adaptation in Vision-Language Models
- Game-TARS: Pretrained Foundation Models for Scalable Generalist Multimodal Game Agents
- Look and Tell: A Dataset for Multimodal Grounding Across Egocentric and Exocentric Views
- Jarvis: Towards Personalized AI Assistant via Personal KV-Cache Retrieval
- Can They Dixit? Yes they Can! Dixit as a Playground for Multimodal Language Model Capabilities
- [De|Re]constructing VLMs' Reasoning in Counting
- Unified Reinforcement and Imitation Learning for Vision-Language Models
- Semantic World Models
- Is This Tracker On? A Benchmark Protocol for Dynamic Tracking
- PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
- Diagnosing Bottlenecks in Data Visualization Understanding by Vision-Language Models
- GOPLA: Generalizable Object Placement Learning via Synthetic Augmentation of Human Arrangement
- InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
- NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
- DeepMMSearch-R1: Empowering Multimodal LLMs in Multimodal Web Search
- Detect Anything via Next Point Prediction
- MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Scaling Language-Centric Omnimodal Representation Learning
- Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- ViSurf: Visual Supervised-and-Reinforcement Fine-Tuning for Large Vision-and-Language Models
- MDSEval: A Meta-Evaluation Benchmark for Multimodal Dialogue Summarization
- How to Teach Large Multimodal Models New Skills
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- BLAZER: Bootstrapping LLM-based Manipulation Agents with Zero-Shot Data Generation
- Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
- From Behavioral Performance to Internal Competence: Interpreting Vision-Language Models with VLM-Lens
- Zero-Shot Fine-Grained Image Classification Using Large Vision-Language Models
- Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
- Efficient Test-Time Scaling for Small Vision-Language Models
- VIRTUE: Visual-Interactive Text-Image Universal Embedder
- DeepSketcher: Internalizing Visual Manipulation for Multimodal Reasoning
- Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
- VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
- Visual serial processing deficits explain divergences in human and VLM reasoning
- Skip-It? Theoretical Conditions for Layer Skipping in Vision-Language Models
- GUI-PRA: Process Reward Agent for GUI Tasks
- Color Names in Vision-Language Models
- Rule-Based Reinforcement Learning for Document Image Classification with Vision Language Models
- Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- ArchGPT: Understanding the World's Architectures with Large Multimodal Models
- Cross-Modal Instructions for Robot Motion Generation
- Queryable 3D Scene Representation: A Multi-Modal Framework for Semantic Reasoning and Robotic Task Planning
- Bias in the Picture: Benchmarking VLMs with Social-Cue News Images and LLM-as-Judge Assessment
- Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
- OmniVLA: An Omni-Modal Vision-Language-Action Model for Robot Navigation
- OmniBridge: Unified Multimodal Understanding, Generation, and Retrieval via Latent Space Alignment
- A Shortcut-aware Video-QA Benchmark for Physical Understanding via Minimal Video Pairs
- BaseReward: A Strong Baseline for Multimodal Reward Model
- ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Decoupled Proxy Alignment: Mitigating Language Prior Conflict for Multimodal Alignment in MLLM
- AToken: A Unified Tokenizer for Vision
- SAIL-VL2 Technical Report
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- Lost in Embeddings: Information Loss in Vision-Language Models
- MindVL: Towards Efficient and Effective Training of Multimodal Large Language Models on Ascend NPUs
- Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations
- Towards Understanding Visual Grounding in Visual Language Models
- Towards Reliable and Interpretable Document Question Answering via VLMs
- Measuring Epistemic Humility in Multimodal Large Language Models
- Towards Better Dental AI: A Multimodal Benchmark and Instruction Dataset for Panoramic X-ray Analysis
- WebMMU: A Benchmark for Multimodal Multilingual Website Understanding and Code Generation
- ToM-SSI: Evaluating Theory of Mind in Situated Social Interactions
- Long-Horizon Visual Imitation Learning via Plan and Code Reflection
- Weakly-Supervised Learning of Dense Functional Correspondences
- MoPEQ: Mixture of Mixed Precision Quantized Experts
- UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning
- LExI: Layer-Adaptive Active Experts for Efficient MoE Model Inference
- Reinforced Visual Perception with Tools
- Improving Large Vision and Language Models by Learning from a Panel of Peers
- Kwai Keye-VL 1.5 Technical Report
- Robix: A Unified Model for Robot Interaction, Reasoning and Planning
- Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
- EO-1: An Open Unified Embodied Foundation Model for General Robot Control
- InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
- Survey of Vision-Language-Action Models for Embodied Manipulation
- Embodied-R1: Reinforced Embodied Reasoning for General Robotic Manipulation
- LangVision-LoRA-NAS: Neural Architecture Search for Variable LoRA Rank in Vision Language Models
- BigCharts-R1: Enhanced Chart Reasoning with Visual Reinforcement Finetuning
- SwarmVLM: VLM-Guided Impedance Control for Autonomous Navigation of Heterogeneous Robots in Dynamic Warehousing
- MolmoAct: Action Reasoning Models that can Reason in Space
- HOLODECK 2.0: Vision-Language-Guided 3D World Generation with Editing
- Open Scene Graphs for Open-World Object-Goal Navigation
- Point2Act: Efficient 3D Distillation of Multimodal LLMs for Zero-Shot Context-Aware Grasping
- Evaluating Variance in Visual Question Answering Benchmarks
- InspectVLM: Unified in Theory, Unreliable in Practice
- Toward Scalable Video Narration: A Training-free Approach Using Multimodal Large Language Models
- UniLiP: Adapting CLIP for Unified Multimodal Understanding, Generation and Editing
- Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos
- Mining Contextualized Visual Associations from Images for Creativity Understanding
- When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
- Describe Anything Model for Visual Question Answering on Text-rich Images
- Hyperphantasia: A Benchmark for Evaluating the Mental Visualization Capabilities of Multimodal LLMs
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
- MapIQ: Evaluating Multimodal Large Language Models for Map Question Answering
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
- If open source is to win, it must go public
- Unlocking Speech Instruction Data Potential with Query Rewriting
- VIP: Visual Information Protection through Adversarial Attacks on Vision-Language Models
- ViLU: Learning Vision-Language Uncertainties for Failure Prediction
- Corvid: Improving Multimodal Large Language Models Towards Chain-of-Thought Reasoning
- Impact of Pretraining Word Co-occurrence on Compositional Generalization in Multimodal Models
- Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
- 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds
- NeoBabel: A Multilingual Open Tower for Visual Generation
- High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning
- Text-Guided Token Communication for Wireless Image Transmission
- VERITAS: Verification and Explanation of Realness in Images for Transparency in AI Systems
- Spatio-Temporal LLM: Reasoning about Environments and Actions
- Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- Vision-Language Models Can't See the Obvious
- Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
- EgoEverything: A Benchmark for Human Behavior Inspired Long Context Egocentric Video Understanding in AR Environment
- VLMs have Tunnel Vision: Evaluating Nonlocal Visual Reasoning in Leading VLMs
- cVLA: Towards Efficient Camera-Space VLAs
- Kwai Keye-VL Technical Report
- SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
- Long-Tailed Distribution-Aware Router For Mixture-of-Experts in Large Vision-Language Model
- Following the Clues: Experiments on Person Re-ID using Cross-Modal Intelligence
- VoyagerVision: Investigating the Role of Multi-modal Information for Open-ended Learning Systems
- StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
- Scaffolding Dexterous Manipulation with Vision-Language Models
- Unblocking Fine-Grained Evaluation of Detailed Captions: An Explaining AutoRater and Critic-and-Revise Pipeline
- MedTVT-R1: A Multimodal LLM Empowering Medical Reasoning and Diagnosis
- OpenEvents V1: Large-Scale Benchmark Dataset for Multimodal Event Grounding
- Adapting Vision-Language Models for Evaluating World Models
- DRAMA-X: A Fine-grained Intent Prediction and Risk Reasoning Benchmark For Driving
- Open World Scene Graph Generation using Vision Language Models
- Visual symbolic mechanisms: Emergent symbol processing in vision language models
- GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
- MoTE: Mixture of Ternary Experts for Memory-efficient Large Multimodal Models
- VL-GenRM: Enhancing Vision-Language Verification via Vision Experts and Iterative Training
- Prompting with the Future: Open-World Model Predictive Control with Interactive Digital Twins
- Touch begins where vision ends: Generalizable policies for contact-rich manipulation
- Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
- MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping
- Movie Facts and Fibs (MF2): A Benchmark for Long Movie Understanding
- VideoMolmo: Spatio-Temporal Grounding Meets Pointing
- CIVET: Systematic Evaluation of Understanding in VLMs
- When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
- OpenVision: A Fully-Open, Cost-Effective Family of Advanced Vision Encoders for Multimodal Learning
- Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning
- MiMo-VL Technical Report
- Sign Language: Towards Sign Understanding for Robot Autonomy
- From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
- HoMeR: Learning In-the-Wild Mobile Manipulation via Hybrid Imitation and Whole-Body Control
- Common Inpainted Objects In-N-Out of Context
- RoboOS: A Hierarchical Embodied Framework for Cross-Embodiment and Multi-Agent Collaboration
- Time Blindness: Why Video-Language Models Can't See What Humans Can?
- Grounded Reinforcement Learning for Visual Reasoning
- Knowledge Insulating Vision-Language-Action Models: Train Fast, Run Fast, Generalize Better
- FAMA: The First Large-Scale Open-Science Speech Foundation Model for English and Italian
- HapticVLM: VLM-Driven Texture Recognition Aimed at Intelligent Haptic Interaction
- Investigating Mechanisms for In-Context Vision Language Binding
- NegVQA: Can Vision Language Models Understand Negation?
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs
- TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
- ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models
- VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
- Large Language Models for Planning: A Comprehensive and Systematic Survey
- Robot Operation of Home Appliances by Reading User Manuals
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
- HAND Me the Data: Fast Robot Adaptation via Hand Path Retrieval
- FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
- R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
- CoRI: Communication of Robot Intent for Physical Human-Robot Interaction
- GRE Suite: Geo-localization Inference via Fine-Tuned Vision-Language Models and Enhanced Reasoning Chains
- Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
- Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
- Panoptic Captioning: An Equivalence Bridge for Image and Text
- Blind Spot Navigation: Evolutionary Discovery of Sensitive Semantic Concepts for LVLMs
- UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
- Efficient Agent Training for Computer Use
- Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
- GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
- ViPlan: A Benchmark for Visual Planning with Symbolic Predicates and Vision-Language Models
- VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
- Real-Time Out-of-Distribution Failure Prevention via Multi-Modal Reasoning
- PointArena: Probing Multimodal Grounding Through Language-Guided Pointing
- Task-Core Memory Management and Consolidation for Long-term Continual Learning
- Behind Maya: Building a Multilingual Vision Language Model
- SAMChat: Introducing Chain of Thought Reasoning and GRPO to a Multimodal Small Language Model for Small Scale Remote Sensing
- Selftok: Discrete Visual Tokens of Autoregression, by Diffusion, and for Reasoning
- 3D CAVLA: Leveraging Depth and 3D Context to Generalize Vision Language Action Models for Unseen Tasks
- Leveraging Vision-Language Models for Visual Grounding and Analysis of Automotive UI
- Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA
- RESAnything: Attribute Prompting for Arbitrary Referring Segmentation
- Kwai Keye-VL-2.0 Technical Report
- Would you still call this Dax? Novel Visual References in VLMs and Humans
- FINER: MLLMs Hallucinate under Fine-grained Negative Queries
- Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models
- Video-Based Reward Modeling for Computer-Use Agents
- LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
- SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data
- Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models
- Anthropogenic Regional Adaptation in Multimodal Vision-Language Model
- Can Vision-Language Models Solve the Shell Game?
- Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
- Contact-Anchored Policies: Contact Conditioning Creates Strong Robot Utility Models
- Unmasking the Illusion of Embodied Reasoning in Vision-Language-Action Models
- HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
- Beyond Perception Errors: Semantic Fixation in Large Vision-Language Models
- MixAtlas: Uncertainty-aware Data Mixture Optimization for Multimodal LLM Midtraining
- Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
- Robotic Task Ambiguity Resolution via Natural Language Interaction
- DyMU: Dynamic Merging and Virtual Unmerging for Efficient VLMs
- VideoVista-CulturalLingo: 360^∘ Horizons-Bridging Cultures, Languages, and Domains in Video Comprehension
- π0.5: a Vision-Language-Action Model with Open-World Generalization
- CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
- Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
- ESPLoRA: Enhanced Spatial Precision with Low-Rank Adaption in Text-to-Image Diffusion Models for High-Definition Synthesis
- ChartQA-X: Generating Explanations for Visual Chart Reasoning
- HoloCount: A Holistic Visual Counting Benchmark for MLLMs
- Benchmarking Vision Language Models on German Factual Data
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
- Perception-R1: Pioneering Perception Policy with Reinforcement Learning
- Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
- SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data
- SmolVLM: Redefining small and efficient multimodal models
Discussions
- Really appreciating the effort of allenai with Molmo&Pixmo to release a fully open VLM along with its training data, great work! Openness is key arxiv.org/abs/2409.17146 [bsky, 8 points, 0 comments]
- Finally, we updated our technical report, including in-depth implementation details and extensive ablation experiments revealing the important model and data design decisions! arxiv.org/abs/2409.17146 [bsky, 7 points, 1 comments]
- oh no doubt. OOD generalization is possible, just need to specifically create a lot of synthetic data to induce learning generalizable representations of those features in a model.
my kingdom for 10 [bsky, 2 points, 0 comments]
- 1/ 🤖 Molmo's data approach: let humans speak freely for 60s about an image. Richer, more natural annotations than structured prompts. @benno_krojer appreciates the design.
https://x.com/benno_krojer/ [bsky, 0 points, 1 comments]
Related