2019/08/30 by Taehee Jung, Dongyeop Kang, Jung, Taehee +5 · 2 citations
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1908.11723
openalex publication_date 2019/08/30 · openalex created_date 2022/07/28 · openalex updated_date 2026/07/28
Despite the recent developments on neural summarization systems, the\nunderlying logic behind the improvements from the systems and its\ncorpus-dependency remains largely unexplored. Position of sentences in the\noriginal text, for example, is a well known bias for news summarization.\nFollowing in the spirit of the claim that summarization is a combination of\nsub-functions, we define three sub-aspects of summarization: position,\nimportance, and diversity and conduct an extensive analysis of the biases of\neach sub-aspect with respect to the domain of nine different summarization\ncorpora (e.g., news, academic papers, meeting minutes, movie script, books,\nposts). We find that while position exhibits substantial bias in news articles,\nthis is not the case, for example, with academic papers and meeting minutes.\nFurthermore, our empirical study shows that different types of summarization\nsystems (e.g., neural-based) are composed of different degrees of the\nsub-aspects. Our study provides useful lessons regarding consideration of\nunderlying sub-aspects when collecting a new summarization dataset or\ndeveloping a new system.\n