UNITER: UNiversal Image-TExt Representation Learning
2020/01/01 by Yen-Chun Chen, Linjie Li, Licheng Yu +5 · 104 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications
paper · doi:10.1007/978-3-030-58577-8_7
crossref issued 2020/01/01 · crossref published 2020/01/01 · crossref published-print 2020/01/01 · openalex publication_date 2020/01/01 · crossref created 2020/09/23 · crossref published-online 2020/09/24 · crossref deposited 2024/09/23 · openalex created_date 2025/10/10 · crossref indexed 2026/07/29 · openalex updated_date 2026/07/31
Cited by
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- Case Relation Transformer: A Crossmodal Language Generation Model for Fetching Instructions
- Detecting Hate Speech in Multi-modal Memes
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- What Vision-Language Models `See' when they See Scenes
- LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- Causal Debiasing for Visual Commonsense Reasoning
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- Graph4MM: Weaving Multimodal Learning with Structural Information
- Exploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- A Multimodal Approach to Heritage Preservation in the Context of Climate Change
- When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- Towards General Purpose Vision Systems
- Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Vision Language Models: A Survey of 26K Papers
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
- Concept Retrieval -- What and How?
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Probing Image-Language Transformers for Verb Understanding
- A novel fusion architecture for detecting Parkinson’s Disease using semi-supervised speech embeddings
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Multilingual Vision-Language Models, A Survey
- Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- A multimodal approach to heritage preservation in the context of climate change
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Data Leakage in Visual Datasets
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Text-Based Person Search with Limited Data
- Understanding the Role of Scene Graphs in Visual Question Answering
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Towards Understanding Visual Grounding in Visual Language Models
- Scaling Up Vision-Language Pre-training for Image Captioning
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Separating Skills and Concepts for Novel Visual Question Answering
- A Recurrent Vision-and-Language BERT for Navigation
- Florence: A New Foundation Model for Computer Vision
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Language Grounding with 3D Objects
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Multimodal Conditionality for Natural Language Generation
- Improving Joint Learning of Chest X-Ray and Radiology Report by Word Region Alignment
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Agentic Design Review System
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Natural Language-Driven Viewpoint Navigation for Volume Exploration via Semantic Block Representation
- Adversarial Video Promotion Against Text-to-Video Retrieval
- RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
- MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- Closing the Modality Gap for Mixed Modality Search
- Open-set Cross Modal Generalization via Multimodal Unified Representation
- PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling
- M6-T: Exploring Sparse Expert Models and Beyond
- Check It Again: Progressive Visual Question Answering via Visual Entailment
- End-to-end Multi-modal Video Temporal Grounding
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
- Caption Enriched Samples for Improving Hateful Memes Detection
- A Multimodal Sentiment Dataset for Video Recommendation
- Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining
- JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs
- CIGLI: Conditional Image Generation from Language & Image
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Smart Routing for Multimodal Video Retrieval: When to Search What
- Scene-Intuitive Agent for Remote Embodied Visual Grounding