2020/03/06 by M. Tanaka, Mikihiro Tanaka, Tanaka, Mikihiro +2 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial intelligence #Closed captioning #Computer Vision and Pattern Recognition (cs.CV) #Computer graphics (images) #Computer science #FOS: Computer and information sciences #Image (mathematics) #Linguistics #Multimodal Machine Learning Applications #Natural language processing #Video Analysis and Summarization #Vocabulary #cs.CV
paper · pdf · doi:10.48550/arxiv.2003.03305
published in arXiv (Cornell University) (Cornell University)
arxiv created 2020/03/06 · openalex publication_date 2020/03/06 · arxiv updated 2020/03/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
In this study, we introduce a low cost method for generating descriptions from images containing novel objects. Generally, constructing a model, which can explain images with novel objects, is costly because of the following: (1) collecting a large amount of data for each category, and (2) retraining the entire system. If humans see a small number of novel objects, they are able to estimate their properties by associating their appearance with known objects. Accordingly, we propose a method that can explain images with novel objects without retraining using the word embeddings of the objects estimated from only a small number of image features of the objects. The method can be integrated with general image-captioning models. The experimental results show the effectiveness of our approach.