MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding
2025/11/16 by Zhanheng Nie, Nie, Zhanheng, Chenghan Fu +13 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #cs.AI #cs.CV #cs.IR #cs.LG
paper · pdf · doi:10.48550/arxiv.2511.12449
Accepted by the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2026. 11 pages, 7 figures
arxiv created 2026/07/29 · arxiv updated 2026/07/31
Abstract
Recent Multimodal Large Language Models (MLLMs) have significantly advanced e-commerce product understanding. However, they still face three challenges: (i) the modality imbalance induced by modality mixed training; (ii) underutilization of the intrinsic alignment relationships among visual and textual information within a product; and (iii) limited handling of noise in e-commerce multimodal data. To address these, we propose MOON2.0, a dynamic modality-balanced MultimOdal representation learning framework for e-commerce prOduct uNderstanding. It comprises: (1) a Modality-driven Mixture-of-Experts (MoE) that adaptively processes input samples by their modality composition, enabling Multimodal Joint Learning to mitigate the modality imbalance; (2) a Dual-level Alignment method to better leverage semantic alignment properties inside individual products; and (3) an MLLM-based Image-text Co-augmentation strategy that integrates textual enrichment with visual expansion, coupled with Dynamic Sample Filtering to improve training data quality. We further release MBE2.0, a co-augmented Multimodal representation Benchmark for E-commerce representation learning and evaluation at https://huggingface.co/datasets/ZHNie/MBE2.0. Experiments show that MOON2.0 delivers state-of-the-art zero-shot performance on MBE2.0 and multiple public datasets. Furthermore, attention-based heatmap visualization provides qualitative evidence of improved multimodal alignment of MOON2.0.
Citations
- MOON Embedding: Multimodal Representation Learning for E-commerce Search Advertising
- EcomMMMU: Strategic Utilization of Visuals for Robust Multimodal E-commerce Models
- UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
- MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding
- InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Qwen2.5-VL Technical Report
- MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling
- GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
- MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval
- MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs
- Unified Generative and Discriminative Training for Multi-modal Large Language Models
- Captions Speak Louder than Images: Generalizing Foundation Models for E-commerce from High-quality Multimodal Instruction Data
- VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
- MRSE: An Efficient Multi-modality Retrieval System for Large Scale E-commerce
- Modality-Balanced Learning for Multimedia Recommendation
- LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
- Multimodal Pretraining, Adaptation, and Generation for Recommendation: A Survey
- eCeLLM: Generalizing Large Language Models for E-commerce from Large-scale, High-quality Instruction Data
- Cross-Domain Product Representation Learning for Rich-Content E-Commerce
- LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following
- StableRep: Synthetic Images from Text-to-Image Models Make Strong Visual Representation Learners
- Learning Instance-Level Representation for Large-Scale Multi-Modal Pretraining in E-commerce
- CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion
- GPT-4 Technical Report
- PMR: Prototypical Modal Rebalance for Multimodal Learning
- Flamingo: a Visual Language Model for Few-Shot Learning
- Contrastive language and vision learning of general fashion concepts
- CommerceMM: Large-Scale Commerce MultiModal Representation Learning with Omni Retrieval
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
- Florence: A New Foundation Model for Computer Vision
- M5Product: Self-harmonized Contrastive Learning for E-commercial Multi-modal Pretraining
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining
- Align before Fuse: Vision and Language Representation Learning with Momentum Distillation
- Embedding-based Product Retrieval in Taobao Search
- Learning Transferable Visual Models From Natural Language Supervision
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- FashionBERT: Text and Image Matching with Adaptive Loss for Cross-modal Retrieval
- A Multimodal Recommender System for Large-scale Assortment Generation in E-commerce
- Automatic Spatially-aware Fashion Concept Discovery
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Cited by
Related