UNITER: UNiversal Image-TExt Representation Learning
2020/01/01 by Yen-Chun Chen, Linjie Li, Licheng Yu +5 · 180 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning #Multimodal Machine Learning Applications
paper · doi:10.1007/978-3-030-58577-8_7
crossref issued 2020/01/01 · crossref published 2020/01/01 · crossref published-print 2020/01/01 · openalex publication_date 2020/01/01 · crossref created 2020/09/23 · crossref published-online 2020/09/24 · crossref deposited 2024/09/23 · openalex created_date 2025/10/10 · crossref indexed 2026/07/29 · openalex updated_date 2026/07/31
Cited by
- X-GGM: Graph Generative Modeling for Out-of-Distribution Generalization in Visual Question Answering
- CRIC: A VQA Dataset for Compositional Reasoning on Vision and Commonsense
- Case Relation Transformer: A Crossmodal Language Generation Model for Fetching Instructions
- Detecting Hate Speech in Multi-modal Memes
- Probing Inter-modality: Visual Parsing with Self-Attention for Vision-Language Pre-training
- What Vision-Language Models `See' when they See Scenes
- LayoutXLM: Multimodal Pre-training for Multilingual Visually-rich Document Understanding
- Vision and Structured-Language Pretraining for Cross-Modal Food Retrieval
- Referring Transformer: A One-step Approach to Multi-task Visual Grounding
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- Causal Debiasing for Visual Commonsense Reasoning
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- iReason: Multimodal Commonsense Reasoning using Videos and Natural Language with Interpretability
- Graph4MM: Weaving Multimodal Learning with Structural Information
- Exploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction
- CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Spatial Preference Rewarding for MLLMs Spatial Understanding
- A Multimodal Approach to Heritage Preservation in the Context of Climate Change
- When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- Towards General Purpose Vision Systems
- Fall into a Pit, Gain in a Wit: Cognitive-Guided Harmful Meme Detection via Misjudgment Risk Pattern Retrieval
- Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Vision Language Models: A Survey of 26K Papers
- Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering
- Concept Retrieval -- What and How?
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs
- MultiFair: Multimodal Balanced Fairness-Aware Medical Classification with Dual-Level Gradient Modulation
- SETR: A Two-Stage Semantic-Enhanced Framework for Zero-Shot Composed Image Retrieval
- Multimodal Large Language Models Meet Multimodal Emotion Recognition and Reasoning: A Survey
- Probing Image-Language Transformers for Verb Understanding
- A novel fusion architecture for detecting Parkinson’s Disease using semi-supervised speech embeddings
- Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Multilingual Vision-Language Models, A Survey
- Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
- VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding
- CLIP-Adapter: Better Vision-Language Models with Feature Adapters
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- A multimodal approach to heritage preservation in the context of climate change
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Align Where the Words Look: Cross-Attention-Guided Patch Alignment with Contrastive and Transport Regularization for Bengali Captioning
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- Self-Supervised Cross-Modal Learning for Image-to-Point Cloud Registration
- OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
- Data Leakage in Visual Datasets
- Improving Generalized Visual Grounding with Instance-aware Joint Learning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Explain Before You Answer: A Survey on Compositional Visual Reasoning
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Text-Based Person Search with Limited Data
- Understanding the Role of Scene Graphs in Visual Question Answering
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Towards Understanding Visual Grounding in Visual Language Models
- Scaling Up Vision-Language Pre-training for Image Captioning
- Enhancing 3D Medical Image Understanding with Pretraining Aided by 2D Multimodal Large Language Models
- Separating Skills and Concepts for Novel Visual Question Answering
- A Recurrent Vision-and-Language BERT for Navigation
- Florence: A New Foundation Model for Computer Vision
- CPT: Colorful Prompt Tuning for Pre-trained Vision-Language Models
- HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training
- Language Grounding with 3D Objects
- Structure-aware Contrastive Learning for Diagram Understanding of Multimodal Models
- Multimodal Contrastive Training for Visual Representation Learning
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- HCCM: Hierarchical Cross-Granularity Contrastive and Matching Learning for Natural Language-Guided Drones
- Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning
- WenLan: Bridging Vision and Language by Large-Scale Multi-Modal Pre-Training
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- Multimodal Conditionality for Natural Language Generation
- Improving Joint Learning of Chest X-Ray and Radiology Report by Word Region Alignment
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Agentic Design Review System
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- Vision Generalist Model: A Survey
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Natural Language-Driven Viewpoint Navigation for Volume Exploration via Semantic Block Representation
- Adversarial Video Promotion Against Text-to-Video Retrieval
- RegionMed-CLIP: A Region-Aware Multimodal Contrastive Learning Pre-trained Model for Medical Image Understanding
- MultiCheck: Strengthening Web Trust with Unified Multimodal Fact Verification
- ETTA: Efficient Test-Time Adaptation for Vision-Language Models through Dynamic Embedding Updates
- Boosting Visual Knowledge-Intensive Training for LVLMs Through Causality-Driven Visual Object Completion
- Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
- Closing the Modality Gap for Mixed Modality Search
- Open-set Cross Modal Generalization via Multimodal Unified Representation
- PhotoChat: A Human-Human Dialogue Dataset with Photo Sharing Behavior for Joint Image-Text Modeling
- M6-T: Exploring Sparse Expert Models and Beyond
- Check It Again: Progressive Visual Question Answering via Visual Entailment
- End-to-end Multi-modal Video Temporal Grounding
- A Picture May Be Worth a Hundred Words for Visual Question Answering
- LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
- StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
- Caption Enriched Samples for Improving Hateful Memes Detection
- A Multimodal Sentiment Dataset for Video Recommendation
- Systematic Generalization on gSCAN: What is Nearly Solved and What is Next?
- Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models
- Product1M: Towards Weakly Supervised Instance-Level Product Retrieval via Cross-modal Pretraining
- JointGT: Graph-Text Joint Representation Learning for Text Generation from Knowledge Graphs
- CIGLI: Conditional Image Generation from Language & Image
- SimVLM: Simple Visual Language Model Pretraining with Weak Supervision
- Smart Routing for Multimodal Video Retrieval: When to Search What
- Scene-Intuitive Agent for Remote Embodied Visual Grounding
- MS-DPPs: Multi-Source Determinantal Point Processes for Contextual Diversity Refinement of Composite Attributes in Text to Image Retrieval
- A Closer Look at the Robustness of Vision-and-Language Pre-trained Models
- Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
- LVLM-Composer's Explicit Planning for Image Generation
- Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
- Large Language Models for Crash Detection in Video: A Survey of Methods, Datasets, and Challenges
- DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy
- LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs
- MiniVLM: A Smaller and Faster Vision-Language Model
- MoCa: Modality-aware Continual Pre-training Makes Better Bidirectional Multimodal Embeddings
- CrossMap Transformer: A Crossmodal Masked Path Transformer Using Double Back-Translation for Vision-and-Language Navigation
- Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
- MultiBench: Multiscale Benchmarks for Multimodal Representation Learning
- CTAL: Pre-training Cross-modal Transformer for Audio-and-Language Representations
- Mind Your Outliers! Investigating the Negative Impact of Outliers on Active Learning for Visual Question Answering
- LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding
- LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- TACo: Token-aware Cascade Contrastive Learning for Video-Text Alignment
- Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- FREE: Fast and Robust Vision Language Models with Early Exits
- Aligning Multimodal Representations through an Information Bottleneck
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
- Cross-Modal Retrieval Augmentation for Multi-Modal Classification
- CLIPort: What and Where Pathways for Robotic Manipulation
- UNIMO: Towards Unified-Modal Understanding and Generation via Cross-Modal Contrastive Learning
- Iconary: A Pictionary-Based Game for Testing Multimodal Communication with Drawings and Text
- MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping
- Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
- Multimodal Text Style Transfer for Outdoor Vision-and-Language Navigation
- Flexible Tool Selection through Low-dimensional Attribute Alignment of Vision and Language
- On VLMs for Diverse Tasks in Multimodal Meme Classification
- BaryIR: Learning Multi-Source Unified Representation in Continuous Barycenter Space for Generalizable All-in-One Image Restoration
- From Data to Modeling: Fully Open-vocabulary Scene Graph Generation
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- Team RUCAIM3 Technical Report at ActivityNet 2021: Entities Object Localization
- LA-RCS: LLM-Agent-Based Robot Control System
- CHAOS: Chart Analysis with Outlier Samples
- FlightGPT: Towards Generalizable and Interpretable UAV Vision-and-Language Navigation with Vision-Language Models
- LAMeTA: Intent-Aware Agentic Network Optimization via a Large AI Model-Empowered Two-Stage Approach
- Mitigating Hallucinations via Inter-Layer Consistency Aggregation in Large Vision-Language Models
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- DPSeg: Dual-Prompt Cost Volume Learning for Open-Vocabulary Semantic Segmentation
- VIGIL: Vision-Language Guided Multiple Instance Learning Framework for Ulcerative Colitis Histological Healing Prediction
- Survey of Visual-Semantic Embedding Methods for Zero-Shot Image Retrieval
- A Transformer-based Cross-modal Fusion Model with Adversarial Training for VQA Challenge 2021
- ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
- Compositional Image-Text Matching and Retrieval by Grounding Entities
- PREMISE: Matching-based Prediction for Accurate Review Recommendation
- VIOLET : End-to-End Video-Language Transformers with Masked Visual-token Modeling
- CUPID: Adaptive Curation of Pre-training Data for Video-and-Language Representation Learning
- Generative Engine Optimization: A VLM and Agent Framework for Pinterest Acquisition Growth
- UFO: A UniFied TransfOrmer for Vision-Language Representation Learning
- M6: A Chinese Multimodal Pretrainer
- TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
- Investigating the Effect of Parallel Data in the Cross-Lingual Transfer for Vision-Language Encoders
- A Survey of Task-Oriented Knowledge Graph Reasoning: Status, Applications, and Prospects
- Achieving Human Parity on Visual Question Answering
- Multimodal Plant Root Phenotyping with Integration of 3D Skeleton Extraction and Language Analysis
- Multimodal graph representation learning for website generation based on visual sketch
- Cracking the Code of Action: a Generative Approach to Affordances for Reinforcement Learning
- Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
- K2MUSE: A human lower limb multimodal dataset under diverse conditions for facilitating rehabilitation robotics
- Data Efficient Masked Language Modeling for Vision and Language
- Neuro-Symbolic Representations for Video Captioning: A Case for Leveraging Inductive Biases for Vision and Language
- VLLFL: A Vision-Language Model Based Lightweight Federated Learning Framework for Smart Agriculture
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- AGO: Adaptive Grounding for Open World 3D Occupancy Prediction
- Graph Learning-Driven Multi-Vessel Association: Fusing Multimodal Data for Maritime Intelligence
- A Lightweight Large Vision-language Model for Multimodal Medical Images
- SCRAMBLe : Enhancing Multimodal LLM Compositionality with Synthetic Preference Data