2015/09/07 by Cheng-Tao Chung, Wei-Ning Hsu, Chung, Cheng-Tao +7
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Music and Audio Processing #Speech Recognition and Synthesis #Speech and Audio Processing #cs.CL
paper · pdf · doi:10.48550/arxiv.1509.02217
Accepted by ICASSP 2015
arxiv created 2015/09/07 · openalex publication_date 2015/09/07 · arxiv updated 2015/09/09 · openalex created_date 2022/10/01 · openalex updated_date 2026/07/28
This paper presents a novel approach for enhancing the multiple sets of acoustic patterns automatically discovered from a given corpus. In a previous work it was proposed that different HMM configurations (number of states per model, number of distinct models) for the acoustic patterns form a two-dimensional space. Multiple sets of acoustic patterns automatically discovered with the HMM configurations properly located on different points over this two-dimensional space were shown to be complementary to one another, jointly capturing the characteristics of the given corpus. By representing the given corpus as sequences of acoustic patterns on different HMM sets, the pattern indices in these sequences can be relabeled considering the context consistency across the different sequences. Good improvements were observed in preliminary experiments of pattern spoken term detection (STD) performed on both TIMIT and Mandarin Broadcast News with such enhanced patterns.