Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
2020/04/13 by Li, Xiujun, Yin, Xi, Li, Chunyuan +9 · 34 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG)
paper · doi:10.48550/arxiv.2004.06165
Abstract
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be pre-trained and use self-attention to learn image-text semantic alignments in a brute force manner, in this paper, we propose a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments. Our method is motivated by the observation that the salient objects in an image can be accurately detected, and are often mentioned in the paired text. We pre-train an Oscar model on the public corpus of 6.5 million text-image pairs, and fine-tune it on downstream tasks, creating new state-of-the-arts on six well-established vision-language understanding and generation tasks.
Cited by
- Investigating Spatial Attention Bias in Vision-Language Models
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- Language-driven Fine-grained Retrieval
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
- Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Graph4MM: Weaving Multimodal Learning with Structural Information
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- Vision Language Models: A Survey of 26K Papers
- Conditional Representation Learning for Customized Tasks
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
- RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- Embedding Font Impression Word Tags Based on Co-occurrence
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- On the Security and Privacy of Federated Learning: A Survey with Attacks, Defenses, Frameworks, Applications, and Future Directions
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Adversarial Video Promotion Against Text-to-Video Retrieval
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
Related