vix.ing · top · new · best · stats · spec

R3Net:Relation-embedded Representation Reconstruction Network for\n Change Captioning

2021/10/19 by Yunbin Tu, Liang Li, Tu, Yunbin +7
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization

paper · pdf · doi:10.48550/arxiv.2110.10328

openalex publication_date 2021/10/19 · openalex created_date 2021/11/22 · openalex updated_date 2026/07/28

Abstract

Change captioning is to use a natural language sentence to describe the\nfine-grained disagreement between two similar images. Viewpoint change is the\nmost typical distractor in this task, because it changes the scale and location\nof the objects and overwhelms the representation of real change. In this paper,\nwe propose a Relation-embedded Representation Reconstruction Network (R3Net)\nto explicitly distinguish the real change from the large amount of clutter and\nirrelevant changes. Specifically, a relation-embedded module is first devised\nto explore potential changed objects in the large amount of clutter. Then,\nbased on the semantic similarities of corresponding locations in the two\nimages, a representation reconstruction module (RRM) is designed to learn the\nreconstruction representation and further model the difference representation.\nBesides, we introduce a syntactic skeleton predictor (SSP) to enhance the\nsemantic interaction between change localization and caption generation.\nExtensive experiments show that the proposed method achieves the\nstate-of-the-art results on two public datasets.\n

Citations

Related