vix.ing · top · new · best · stats

An Effective Contextual Language Modeling Framework for Speech Summarization with Augmented Features

2020/06/01 by Shi-Yan Weng, Weng, Shi-Yan, Tien-Hong Lo +3 · 2 citations
Computer Science · #Artificial intelligence #Automatic summarization #Benchmark (surveying) #Computation and Language (cs.CL) #Computer science #Encoder #FOS: Computer and information sciences #Language model #Natural Language Processing Techniques #Natural language #Natural language processing #Sentence #Speech Recognition and Synthesis #Speech recognition #Topic Modeling #Transformer #cs.CL

paper · pdf · doi:10.48550/arxiv.2006.01189

published in arXiv (Cornell University) (Cornell University) · Accepted by EUSIPCO 2020

arxiv created 2020/06/01 · openalex publication_date 2020/06/01 · arxiv updated 2020/06/03 · openalex created_date 2020/06/12 · openalex updated_date 2026/08/06

Abstract

Tremendous amounts of multimedia associated with speech information are driving an urgent need to develop efficient and effective automatic summarization methods. To this end, we have seen rapid progress in applying supervised deep neural network-based methods to extractive speech summarization. More recently, the Bidirectional Encoder Representations from Transformers (BERT) model was proposed and has achieved record-breaking success on many natural language processing (NLP) tasks such as question answering and language understanding. In view of this, we in this paper contextualize and enhance the state-of-the-art BERT-based model for speech summarization, while its contributions are at least three-fold. First, we explore the incorporation of confidence scores into sentence representations to see if such an attempt could help alleviate the negative effects caused by imperfect automatic speech recognition (ASR). Secondly, we also augment the sentence embeddings obtained from BERT with extra structural and linguistic features, such as sentence position and inverse document frequency (IDF) statistics. Finally, we validate the effectiveness of our proposed method on a benchmark dataset, in comparison to several classic and celebrated speech summarization methods.

Citations

Related