ProM3E: Probabilistic Masked MultiModal Embedding Model for Ecology
2025/11/04 by Sastry, Srikumar, Khanal, Subash, Dhakal, Aayush +4 · 1 citation
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
paper · doi:10.48550/arxiv.2511.02946
Abstract
We introduce ProM3E, a probabilistic masked multimodal embedding model for any-to-any generation of multimodal representations for ecology. ProM3E is based on masked modality reconstruction in the embedding space, learning to infer missing modalities given a few context modalities. By design, our model supports modality inversion in the embedding space. The probabilistic nature of our model allows us to analyse the feasibility of fusing various modalities for given downstream tasks, essentially learning what to fuse. Using these features of our model, we propose a novel cross-modal retrieval approach that mixes inter-modal and intra-modal similarities to achieve superior performance across all retrieval tasks. We further leverage the hidden representation from our model to perform linear probing tasks and demonstrate the superior representation learning capability of our model. All our code, datasets and model will be released at https://vishu26.github.io/prom3e.
Citations
- EcoWikiRS: Learning Ecological Representation of Satellite Images from Weak Supervision with Species Observations and Wikipedia
- Climplicit: Climatic Implicit Embeddings for Global Ecological Tasks
- MaskSDM with Shapley values to improve flexibility, robustness, and explainability in species distribution modeling
- Towards a Unified Copernicus Foundation Model for Earth Vision
- RANGE: Retrieval Augmented Neural Fields for Multi-Resolution Geo-Embeddings
- AnyTouch: Learning Unified Static-Dynamic Representation across Multiple Visuo-tactile Sensors
- Cross the Gap: Exposing the Intra-modal Misalignment in CLIP via Modality Inversion
- WildSAT: Learning Satellite Image Representations from Wildlife Observations
- AnySat: One Earth Observation Model for Many Resolutions, Scales, and Modalities
- Causal Representation Learning from Multimodal Biomedical Observations
- INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
- TaxaBind: A Unified Embedding Space for Ecological Applications
- Probabilistic Language-Image Pre-Training
- Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery
- PSM: Learning Probabilistic Embeddings for Multi-scale Zero-Shot Soundscape Mapping
- OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
- BioTrove: A Large Curated Image Dataset Enabling AI for Biodiversity
- OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning
- 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
- All in One Framework for Multimodal Re-identification in the Wild
- GEOBIND: Binding Text, Image, and Audio through Satellite Images
- Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
- UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
- Binding Touch to Everything: Learning Unified Multimodal Tactile Representations
- Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment
- 4M: Massively Multimodal Masked Modeling
- OneLLM: One Framework to Align All Modalities with Language
- BioCLIP: A Vision Foundation Model for the Tree of Life
- SatCLIP: Global, General-Purpose Location Embeddings with Satellite Imagery
- ViT-Lens: Towards Omni-modal Representations
- OmniVec: Learning robust representations with cross modal sharing
- BirdSAT: Cross-View Contrastive Masked Autoencoders for Bird Species Classification and Mapping
- Penetrative AI: Making LLMs Comprehend the Physical World
- Extending Multi-modal Contrastive Representations
- Vision Transformers Need Registers
- GeoCLIP: Clip-Inspired Alignment between Locations and Images for Effective Worldwide Geo-localization
- Learning Tri-modal Embeddings for Zero-Shot Soundscape Mapping
- NExT-GPT: Any-to-Any Multimodal LLM
- UnIVAL: Unified Model for Image, Video, Audio and Language Tasks
- Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images
- Spatial Implicit Neural Representations for Global-Scale Species Mapping
- Improved Probabilistic Image-Text Representations
- Connecting Multi-modal Contrastive Representations
- ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
- ImageBind: One Embedding Space To Bind Them All
- CSP: Self-Supervised Contrastive Spatial Pre-Training for Geospatial-Visual Representations
- Multi-modal Variational Autoencoders for normative modelling across multiple imaging modalities
- Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture
- MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
- Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
- Probabilistic Compositional Embeddings for Multimodal Image Retrieval
- Probabilistic Representations for Video Contrastive Learning
- MultiMAE: Multi-modal Multi-task Masked Autoencoders
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- Mind the Gap: Understanding the Modality Gap in Multi-modal Contrastive Representation Learning
- Masked Feature Prediction for Self-Supervised Visual Pre-Training
- SimMIM: A Simple Framework for Masked Image Modeling
- Masked Autoencoders Are Scalable Vision Learners
- AudioCLIP: Extending CLIP to Image, Text and Audio
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Learning Transferable Visual Models From Natural Language Supervision
- Supervised Contrastive Learning
- ALBERT: A Lite BERT for Self-supervised Learning of Language Representations
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Presence-Only Geographical Priors for Fine-Grained Image Classification
- Learning Grid Cells as Vector Representation of Self-Position Coupled with Matrix Representation of Self-Motion
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- Two Effects, One Trigger: On the Modality Gap, Object Bias, and Information Imbalance in Contrastive Vision-Language Models
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Cited by
Related