2017/03/23 by Herman Kamper, Shane Settle, Kamper, Herman +5
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
paper · pdf · doi:10.48550/arxiv.1703.08136
openalex publication_date 2017/03/23 · openalex created_date 2022/10/04 · openalex updated_date 2026/07/28
During language acquisition, infants have the benefit of visual cues to\nground spoken language. Robots similarly have access to audio and visual\nsensors. Recent work has shown that images and spoken captions can be mapped\ninto a meaningful common space, allowing images to be retrieved using speech\nand vice versa. In this setting of images paired with untranscribed spoken\ncaptions, we consider whether computer vision systems can be used to obtain\ntextual labels for the speech. Concretely, we use an image-to-words multi-label\nvisual classifier to tag images with soft textual labels, and then train a\nneural network to map from the speech to these soft targets. We show that the\nresulting speech system is able to predict which words occur in an\nutterance---acting as a spoken bag-of-words classifier---without seeing any\nparallel speech and text. We find that the model often confuses semantically\nrelated words, e.g. "man" and "person", making it even more effective as a\nsemantic keyword spotter.\n