2021/04/19 by Sihao Chen, Fan Zhang, Chen, Sihao +5 · 7 citations
Computer Science · #Advanced Text Analysis Techniques #Artificial intelligence #Automatic summarization #Computer science #Context (archaeology) #Contrast (vision) #Discriminative model #Machine learning #Natural Language Processing Techniques #Natural language processing #Selection (genetic algorithm) #Source text #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2104.09061
published in arXiv (Cornell University) (Cornell University) · NAACL'21
arxiv created 2021/04/19 · openalex publication_date 2021/04/19 · arxiv updated 2021/04/20 · openalex created_date 2021/04/26 · openalex updated_date 2026/08/05
Despite significant progress in neural abstractive summarization, recent studies have shown that the current models are prone to generating summaries that are unfaithful to the original context. To address the issue, we study contrast candidate generation and selection as a model-agnostic post-processing technique to correct the extrinsic hallucinations (i.e. information not present in the source text) in unfaithful summaries. We learn a discriminative correction model by generating alternative candidate summaries where named entities and quantities in the generated summary are replaced with ones with compatible semantic types from the source document. This model is then used to select the best candidate as the final output summary. Our experiments and analysis across a number of neural summarization systems show that our proposed method is effective in identifying and correcting extrinsic hallucinations. We analyze the typical hallucination phenomenon by different types of neural summarization systems, in hope to provide insights for future work on the direction.