2018/07/26 by Bo Dai, Dai, Bo, Deming Ye +3
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Multimodal Machine Learning Applications #cs.CV #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.1807.09958
ECCV 2018, first two authors contribute equally
arxiv created 2018/07/26 · openalex publication_date 2018/07/26 · arxiv updated 2018/08/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
RNNs and their variants have been widely adopted for image captioning. In RNNs, the production of a caption is driven by a sequence of latent states. Existing captioning models usually represent latent states as vectors, taking this practice for granted. We rethink this choice and study an alternative formulation, namely using two-dimensional maps to encode latent states. This is motivated by the curiosity about a question: how the spatial structures in the latent states affect the resultant captions? Our study on MSCOCO and Flickr30k leads to two significant observations. First, the formulation with 2D states is generally more effective in captioning, consistently achieving higher performance with comparable parameter sizes. Second, 2D states preserve spatial locality. Taking advantage of this, we visually reveal the internal dynamics in the process of caption generation, as well as the connections between input visual domain and output linguistic domain.