2022/03/15 by Fandong Meng, Zheng, Duo, Meng, Fandong +13
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Human Pose and Action Recognition
paper · pdf · doi:10.48550/arxiv.2203.08362
Visual dialog has witnessed great progress after introducing various\nvision-oriented goals into the conversation, especially such as GuessWhich and\nGuessWhat, where the only image is visible by either and both of the questioner\nand the answerer, respectively. Researchers explore more on visual dialog tasks\nin such kind of single- or perfectly co-observable visual scene, while somewhat\nneglect the exploration on tasks of non perfectly co-observable visual scene,\nwhere the images accessed by two agents may not be exactly the same, often\noccurred in practice. Although building common ground in non-perfectly\nco-observable visual scene through conversation is significant for advanced\ndialog agents, the lack of such dialog task and corresponding large-scale\ndataset makes it impossible to carry out in-depth research. To break this\nlimitation, we propose an object-referring game in non-perfectly co-observable\nvisual scene, where the goal is to spot the difference between the similar\nvisual scenes through conversing in natural language. The task addresses\nchallenges of the dialog strategy in non-perfectly co-observable visual scene\nand the ability of categorizing objects. Correspondingly, we construct a\nlarge-scale multimodal dataset, named SpotDiff, which contains 87k Virtual\nReality images and 97k dialogs generated by self-play. Finally, we give\nbenchmark models for this task, and conduct extensive experiments to evaluate\nits performance as well as analyze its main challenges.\n