vix.ing · top · new · best · stats · spec

Unity in Diversity: Learning Distributed Heterogeneous Sentence\n Representation for Extractive Summarization

2019/12/25 by Abhishek Singh, Manish Gupta, Singh, Abhishek Kumar +4
Computer Science · #Advanced Text Analysis Techniques #Computation and Language (cs.CL) #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1912.11688

openalex publication_date 2019/12/25 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28

Abstract

Automated multi-document extractive text summarization is a widely studied\nresearch problem in the field of natural language understanding. Such\nextractive mechanisms compute in some form the worthiness of a sentence to be\nincluded into the summary. While the conventional approaches rely on human\ncrafted document-independent features to generate a summary, we develop a\ndata-driven novel summary system called HNet, which exploits the various\nsemantic and compositional aspects latent in a sentence to capture document\nindependent features. The network learns sentence representation in a way that,\nsalient sentences are closer in the vector space than non-salient sentences.\nThis semantic and compositional feature vector is then concatenated with the\ndocument-dependent features for sentence ranking. Experiments on the DUC\nbenchmark datasets (DUC-2001, DUC-2002 and DUC-2004) indicate that our model\nshows significant performance gain of around 1.5-2 points in terms of ROUGE\nscore compared with the state-of-the-art baselines.\n

Citations

Related