2018/04/15 by Reuben Cohn-Gordon, Noah D. Goodman, Cohn-Gordon, Reuben +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1804.05417
openalex publication_date 2018/04/15 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
We combine a neural image captioner with a Rational Speech Acts (RSA) model\nto make a system that is pragmatically informative: its objective is to produce\ncaptions that are not merely true but also distinguish their inputs from\nsimilar images. Previous attempts to combine RSA with neural image captioning\nrequire an inference which normalizes over the entire set of possible\nutterances. This poses a serious problem of efficiency, previously solved by\nsampling a small subset of possible utterances. We instead solve this problem\nby implementing a version of RSA which operates at the level of characters\n("a","b","c"...) during the unrolling of the caption. We find that the\nutterance-level effect of referential captions can be obtained with only\ncharacter-level decisions. Finally, we introduce an automatic method for\ntesting the performance of pragmatic speaker models, and show that our model\noutperforms a non-pragmatic baseline as well as a word-level RSA captioner.\n