vix.ing · top · new · best · stats

Captioning Images with Novel Objects via Online Vocabulary Expansion

2020/03/06 by M. Tanaka, Mikihiro Tanaka, Tanaka, Mikihiro +2 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Closed captioning #Computer Vision and Pattern Recognition (cs.CV) #Computer graphics (images) #Computer science #FOS: Computer and information sciences #Image (mathematics) #Linguistics #Multimodal Machine Learning Applications #Natural language processing #Video Analysis and Summarization #Vocabulary #cs.CV

paper · pdf · doi:10.48550/arxiv.2003.03305

published in arXiv (Cornell University) (Cornell University)

arxiv created 2020/03/06 · openalex publication_date 2020/03/06 · arxiv updated 2020/03/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

In this study, we introduce a low cost method for generating descriptions from images containing novel objects. Generally, constructing a model, which can explain images with novel objects, is costly because of the following: (1) collecting a large amount of data for each category, and (2) retraining the entire system. If humans see a small number of novel objects, they are able to estimate their properties by associating their appearance with known objects. Accordingly, we propose a method that can explain images with novel objects without retraining using the word embeddings of the objects estimated from only a small number of image features of the objects. The method can be integrated with general image-captioning models. The experimental results show the effectiveness of our approach.

Citations

Related