Oscar: Object-Semantics Aligned Pre-training for Vision-Language Tasks
2020/04/13 by Xiujun Li, Xi Yin, Li, Xiujun +21 · 69 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #cs.CL #cs.CV #cs.IR #cs.LG
paper · pdf · doi:10.48550/arxiv.2004.06165
ECCV 2020, Code and pre-trained models are released: https://github.com/microsoft/Oscar
arxiv created 2020/07/26 · arxiv updated 2020/07/28
Abstract
Large-scale pre-training methods of learning cross-modal representations on image-text pairs are becoming popular for vision-language tasks. While existing methods simply concatenate image region features and text features as input to the model to be pre-trained and use self-attention to learn image-text semantic alignments in a brute force manner, in this paper, we propose a new learning method Oscar (Object-Semantics Aligned Pre-training), which uses object tags detected in images as anchor points to significantly ease the learning of alignments. Our method is motivated by the observation that the salient objects in an image can be accurately detected, and are often mentioned in the paired text. We pre-train an Oscar model on the public corpus of 6.5 million text-image pairs, and fine-tune it on downstream tasks, creating new state-of-the-arts on six well-established vision-language understanding and generation tasks.
Cited by
- Investigating Spatial Attention Bias in Vision-Language Models
- Neurosymbolic Inference On Foundation Models For Remote Sensing Text-to-image Retrieval With Complex Queries
- Language-driven Fine-grained Retrieval
- Are Neuro-Inspired Multi-Modal Vision-Language Models Resilient to Membership Inference Privacy Leakage?
- Robust Defense Strategies for Multimodal Contrastive Learning: Efficient Fine-tuning Against Backdoor Attacks
- Spatio-Temporal Data Enhanced Vision-Language Model for Traffic Scene Understanding
- SLIP: Structural-aware Language-Image Pretraining for Vision-Language Alignment
- Enhancing Adversarial Transferability in Visual-Language Pre-training Models via Local Shuffle and Sample-based Attack
- Modest-Align: Data-Efficient Alignment for Vision-Language Models
- See, Think, Act: Online Shopper Behavior Simulation with VLM Agents
- Graph4MM: Weaving Multimodal Learning with Structural Information
- Class Prototypes based Contrastive Learning for Classifying Multi-Label and Fine-Grained Educational Videos
- Towards Self-Refinement of Vision-Language Models with Triangular Consistency
- Vision Language Models: A Survey of 26K Papers
- Conditional Representation Learning for Customized Tasks
- CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning
- Team Xiaomi EV-AD VLA: Caption-Guided Retrieval System for Cross-Modal Drone Navigation -- Technical Report for IROS 2025 RoboSense Challenge Track 4
- The Artificial Intelligence Cognitive Examination: A Survey on the Evolution of Multimodal Evaluation From Recognition to Reasoning
- Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
- RACap: Relation-Aware Prompting for Lightweight Retrieval-Augmented Image Captioning
- MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
- DyKen-Hyena: Dynamic Kernel Generation via Cross-Modal Attention for Multimodal Intent Recognition
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Prototype-Aware Multimodal Alignment for Open-Vocabulary Visual Grounding
- Embedding Font Impression Word Tags Based on Co-occurrence
- SEAM: Semantically Equivalent Across Modalities Benchmark for Vision-Language Models
- On the Security and Privacy of Federated Learning: A Survey with Attacks, Defenses, Frameworks, Applications, and Future Directions
- VELVET-Med: Vision and Efficient Language Pre-training for Volumetric Imaging Tasks in Medicine
- WeatherPrompt: Multi-modality Representation Learning for All-Weather Drone Visual Geo-Localization
- RORPCap: Retrieval-based Objects and Relations Prompt for Image Captioning
- Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges
- Adversarial Video Promotion Against Text-to-Video Retrieval
- Modality-Aware Feature Matching: A Comprehensive Review of Single- and Cross-Modality Techniques
- When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
- Describe Anything Model for Visual Question Answering on Text-rich Images
- Unveiling Effective In-Context Configurations for Image Captioning: An External & Internal Analysis
- Computed Tomography Visual Question Answering with Cross-modal Feature Graphing
- Towards Universal & Efficient Model Compression via Exponential Torque Pruning
- Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search
- PEVLM: Parallel Encoding for Vision-Language Models
- Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding
- From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
- VLN-R1: Vision-Language Navigation via Reinforcement Fine-Tuning
- Stepping Out of Similar Semantic Space for Open-Vocabulary Segmentation
- Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs
- FREE: Fast and Robust Vision Language Models with Early Exits
- DiffCAP: Diffusion-based Cumulative Adversarial Purification for Vision Language Models
- Joint Generalized Cosine Similarity: A Novel Method for N-Modal Semantic Alignment Based on Contrastive Learning
- Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model
- On VLMs for Diverse Tasks in Multimodal Meme Classification
- Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
- GeoMM: On Geodesic Perspective for Multi-modal Learning
- Incorporating brain-inspired mechanisms for multimodal learning in artificial intelligence
- Structural-Temporal Coupling Anomaly Detection with Dynamic Graph Transformer
- Compositional Image-Text Matching and Retrieval by Grounding Entities
- Evolution of Accuracy and Visual-Cognitive Errors in a Decade of Vision-Language AI Models
- Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI
- Open World Knowledge Aided Single-Cell Foundation Model with Robust Cross-Modal Cell-Language Pre-training
- Symbolic Representation for Any-to-Any Generative Tasks
- Towards Explainable AI: Multi-Modal Transformer for Video-based Image Description Generation
- When Modalities Fail to Tango: Conformal Backdoor Detection in Multimodal Contrastive Learning
- Zero-Shot, But at What Cost? Unveiling the Hidden Overhead of MILS's LLM-CLIP Framework for Image Captioning
- Enabling Collaborative Parametric Knowledge Calibration for Retrieval-Augmented Vision Question Answering
- TADACap: Time-series Adaptive Domain-Aware Captioning
- COUNTS: Benchmarking Object Detectors and Multimodal Large Language Models under Distribution Shifts
- The Mirage of Performance Gains: Why Contrastive Decoding Fails to Mitigate Object Hallucinations in MLLMs?
- Bayesian Cross-Modal Alignment Learning for Few-Shot Out-of-Distribution Generalization
- A Survey on Efficient Vision-Language Models
- How Can Objects Help Video-Language Understanding?
Related