2015/04/13 by Xiaodong He, Rupesh K. Srivastava, He, Xiaodong +5
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1504.03083
openalex publication_date 2015/04/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
This technical report provides extra details of the deep multimodal similarity model (DMSM) which was proposed in (Fang et al. 2015, arXiv:1411.4952). The model is trained via maximizing global semantic similarity between images and their captions in natural language using the public Microsoft COCO database, which consists of a large set of images and their corresponding captions. The learned representations attempt to capture the combination of various visual concepts and cues.