2013/05/01 by Li Deng, Geoffrey E. Hinton, Brian Kingsbury · 7 citations
Computer Science · #Speech Recognition and Synthesis #Speech and Audio Processing #Music and Audio Processing #Computer science #Artificial neural network #Deep learning #Artificial intelligence #Speech recognition
paper · doi:10.1109/icassp.2013.6639344
openalex publication_date 2013/05/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
In this paper, we provide an overview of the invited and contributed papers presented at the special session at ICASSP-2013, entitled “New Types of Deep Neural Network Learning for Speech Recognition and Related Applications,” as organized by the authors. We also describe the historical context in which acoustic models based on deep neural networks have been developed. The technical overview of the papers presented in our special session is organized into five ways of improving deep learning methods: (1) better optimization; (2) better types of neural activation function and better network architectures; (3) better ways to determine the myriad hyper-parameters of deep neural networks; (4) more appropriate ways to preprocess speech for deep neural networks; and (5) ways of leveraging multiple languages or dialects that are more easily achieved with deep neural networks than with Gaussian mixture models.