2015/10/07 by David Schlangen, Sina Zarrieß, Schlangen, David +3 · 1 citation
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Multimodal Machine Learning Applications
paper · pdf · doi:10.48550/arxiv.1510.02125
openalex publication_date 2015/10/07 · openalex created_date 2022/10/05 · openalex updated_date 2026/07/28
A common use of language is to refer to visually present objects. Modelling\nit in computers requires modelling the link between language and perception.\nThe "words as classifiers" model of grounded semantics views words as\nclassifiers of perceptual contexts, and composes the meaning of a phrase\nthrough composition of the denotations of its component words. It was recently\nshown to perform well in a game-playing scenario with a small number of object\ntypes. We apply it to two large sets of real-world photographs that contain a\nmuch larger variety of types and for which referring expressions are available.\nUsing a pre-trained convolutional neural network to extract image features, and\naugmenting these with in-picture positional information, we show that the model\nachieves performance competitive with the state of the art in a reference\nresolution task (given expression, find bounding box of its referent), while,\nas we argue, being conceptually simpler and more flexible.\n