LXMERT: Learning Cross-Modality Encoder Representations from Transformers
2019/08/20 by Hao Tan, Mohit Bansal, Tan, Hao +1 · 101 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1908.07490
openalex publication_date 2019/08/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Abstract
Vision-and-language reasoning requires an understanding of visual concepts, language semantics, and, most importantly, the alignment and relationships between these two modalities. We thus propose the LXMERT (Learning Cross-Modality Encoder Representations from Transformers) framework to learn these vision-and-language connections. In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder. Next, to endow our model with the capability of connecting vision and language semantics, we pre-train the model with large amounts of image-and-sentence pairs, via five diverse representative pre-training tasks: masked language modeling, masked object prediction (feature regression and label classification), cross-modality matching, and image question answering. These tasks help in learning both intra-modality and cross-modality relationships. After fine-tuning from our pre-trained parameters, our model achieves the state-of-the-art results on two visual question answering datasets (i.e., VQA and GQA). We also show the generalizability of our pre-trained cross-modality model by adapting it to a challenging visual-reasoning task, NLVR2, and improve the previous best result by 22% absolute (54% to 76%). Lastly, we demonstrate detailed ablation studies to prove that both our novel model components and pre-training strategies significantly contribute to our strong results; and also present several attention visualizations for the different encoders. Code and pre-trained models publicly available at: https://github.com/airsplay/lxmert
Citations
Cited by
- VL-RouterBench: A Benchmark for Vision-Language Model Routing
- DDVT: Dynamic Dual-level Vision Transformer Fusion Network for Answer Grounding in Visual Question Answering
- Towards Dual-Brain Minimal Sufficient Representation for Vision-Language Navigation
- ETP-R1: Evolving Topological Planning with Reinforcement Fine-tuning for Vision-Language Navigation in Continuous Environments
- VLNVerse: A Benchmark for Vision-Language Navigation with Versatile, Embodied, Realistic Simulation and Evaluation
- CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
- Semantic Mismatch and Perceptual Degradation: A New Perspective on Image Editing Immunity
- ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- Enhancing Medical Cross-Modal Hashing Retrieval using Dropout-Voting Mixture-of-Experts Fusion
- Language-driven Fine-grained Retrieval
- MIND: Multi-rationale INtegrated Discriminative Reasoning Framework for Multi-modal Large Models
- Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
- Think Before You Drive: World Model-Inspired Multimodal Grounding for Autonomous Vehicles
- Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
- VaMP: Variational Multi-Modal Prompt Learning for Vision-Language Models
- Advanced Data Collection Techniques in Cloud Security: A Multi-Modal Deep Learning Autoencoder Approach
- AnchorOPT: Towards Optimizing Dynamic Anchors for Adaptive Prompt Learning
- Think First, Assign Next (ThiFAN-VQA): A Two-stage Chain-of-Thought Framework for Post-Disaster Damage Assessment
- C2F-Space: Coarse-to-Fine Space Grounding for Spatial Instructions using Vision-Language Models
- Reconstruction-Driven Multimodal Representation Learning for Automated Media Understanding
- Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding
- Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
- Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings
- Thought-For-Food: Reasoning Chain Induced Food Visual Question Answering
- SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
- Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
- Towards Automated Petrography
- FOCUS: Efficient Keyframe Selection for Long Video Understanding
- A Retrospect to Multi-prompt Learning across Vision and Language
- MM-OPERA: Benchmarking Open-ended Association Reasoning for Large Vision-Language Models
- Masked Diffusion Captioning for Visual Feature Learning
- Top-Down Semantic Refinement for Image Captioning
- Gaze-VLM:Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- FlexiReID: Adaptive Mixture of Expert for Multi-Modal Person Re-Identification
- Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures
- Towards a Generalizable Fusion Architecture for Multimodal Object Detection
- Unifying Vision-Language Latents for Zero-label Image Caption Enhancement
- Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
- Multimodal Retrieval-Augmented Generation with Large Language Models for Medical VQA
- Cluster-Aware Prompt Ensemble Learning for Few-Shot Vision-Language Model Adaptation
- Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Spatial-ViLT: Enhancing Visual Spatial Reasoning through Multi-Task Learning
- FusionAdapter for Few-Shot Relation Learning in Multimodal Knowledge Graphs
- Landmark-Guided Knowledge for Vision-and-Language Navigation
- Q-FSRU: Quantum-Augmented Frequency-Spectral For Medical Visual Question Answering
- SynDoc: A Hybrid Discriminative-Generative Framework for Enhancing Synthetic Domain-Adaptive Document Key Information Extraction
- Efficient Self-supervised Vision Transformers for Representation Learning
- Multilingual Vision-Language Models, A Survey
- Resolving Ambiguity in Gaze-Facilitated Visual Assistant Interaction Paradigm
- Spatial Reasoning in Foundation Models: Benchmarking Object-Centric Spatial Understanding
- Integrating Object Interaction Self-Attention and GAN-Based Debiasing for Visual Question Answering
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Human-centric Spatio-Temporal Video Grounding With Visual Transformers
- Cross-Modality Protein Embedding for Compound-Protein Affinity and Contact Prediction
- KVL-BERT: Knowledge Enhanced Visual-and-Linguistic BERT for Visual Commonsense Reasoning
- LCMF: Lightweight Cross-Modality Mambaformer for Embodied Robotics VQA
- Single-Branch Network Architectures to Close the Modality Gap in Multimodal Recommendation
- Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
- M3ET: Efficient Vision-Language Learning for Robotics based on Multimodal Mamba-Enhanced Transformer
- DA-Mamba: Dialogue-aware selective state-space model for multimodal engagement estimation
- How Much Can CLIP Benefit Vision-and-Language Tasks?
- DAFTED: Decoupled Asymmetric Fusion of Tabular and Echocardiographic Data for Cardiac Hypertension Diagnosis
- Copycat vs. Original: Multi-modal Pretraining and Variable Importance in Box-office Prediction
- OpenViDial: A Large-Scale, Open-Domain Dialogue Dataset with Visual Contexts
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- Biomedical Hypothesis Explainability with Graph-Based Context Retrieval
- Understanding the Role of Scene Graphs in Visual Question Answering
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- Towards Understanding Visual Grounding in Visual Language Models
- Attn-Adapter: Attention Is All You Need for Online Few-shot Learner of Vision-Language Model
- EVENT-Retriever: Event-Aware Multimodal Image Retrieval for Realistic Captions
- Attention that does not Explain Away
- Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- ManyModalQA: Modality Disambiguation and QA over Diverse Inputs
- Multimodal Data Storage and Retrieval for Embodied AI: A Survey
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- Q-FSRU: Quantum-Augmented Frequency-Spectral Fusion for Medical Visual Question Answering
- Understanding in Artificial Intelligence
- Reasoning in Computer Vision: Taxonomy, Models, Tasks, and Methodologies
- A Curriculum Learning Approach to Reinforcement Learning: Leveraging RAG for Multimodal Question Answering
- BERT-VQA: Visual Question Answering on Plots
- SemVLP: Vision-Language Pre-training by Aligning Semantics at Multiple Levels
- DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
- Dual-stream Network for Visual Recognition
- AME: Aligned Manifold Entropy for Robust Vision-Language Distillation
- Multimodal attention-based deep learning for Alzheimer’s disease diagnosis
- FLUID: Flow-Latent Unified Integration via Token Distillation for Expert Specialization in Multimodal Learning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Natural Language-Driven Viewpoint Navigation for Volume Exploration via Semantic Block Representation
- Adversarial Video Promotion Against Text-to-Video Retrieval
- Multimodal RAG Enhanced Visual Description
- UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
- VQA support to Arabic Language Learning Educational Tool
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- Traits Run Deep: Enhancing Personality Assessment via Psychology-Guided LLM Representations and Multimodal Apparent Behaviors
- Goal-Based Vision-Language Driving
- Analyzing the Sensitivity of Vision Language Models in Visual Question Answering
- Learning Video Representations using Contrastive Bidirectional Transformer
Related