vix.ing · top · new · best · stats · spec

Incorporating Visual Semantics into Sentence Representations within a\n Grounded Space

2020/02/07 by Patrick Bordes, Éloi Zablocki, Bordes, Patrick +7
Computer Science · Arts and Humanities · Psychology · #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media #Language, Metaphor, and Cognition

paper · pdf · doi:10.48550/arxiv.2002.02734

Abstract

Language grounding is an active field aiming at enriching textual\nrepresentations with visual information. Generally, textual and visual elements\nare embedded in the same representation space, which implicitly assumes a\none-to-one correspondence between modalities. This hypothesis does not hold\nwhen representing words, and becomes problematic when used to learn sentence\nrepresentations --- the focus of this paper --- as a visual scene can be\ndescribed by a wide variety of sentences. To overcome this limitation, we\npropose to transfer visual information to textual representations by learning\nan intermediate representation space: the grounded space. We further propose\ntwo new complementary objectives ensuring that (1) sentences associated with\nthe same visual content are close in the grounded space and (2) similarities\nbetween related elements are preserved across modalities. We show that this\nmodel outperforms the previous state-of-the-art on classification and semantic\nrelatedness tasks.\n

Citations

Cited by

Related