2013/12/21 by Lei Jimmy Ba, Jimmy Ba, Rich Caruana +2 · 4 voices · 1,484 citations
Computer Science · Engineering · #Artificial intelligence #Artificial neural network #Computer science #Convolutional neural network #Deep learning #Deep neural networks #Deep water #Engineering #Generative Adversarial Networks and Image Synthesis #Hidden Markov model #Music and Audio Processing #Speech Recognition and Synthesis #TIMIT #Task (project management) #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1312.6184
published in arXiv (Cornell University) (Cornell University) · final revision coming soon
openalex publication_date 2013/12/21 · arxiv created 2014/10/11 · arxiv updated 2014/10/14 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
Currently, deep neural networks are the state of the art on problems such as speech recognition and computer vision. In this extended abstract, we show that shallow feed-forward networks can learn the complex functions previously learned by deep nets and achieve accuracies previously only achievable with deep models. Moreover, in some cases the shallow neural nets can learn these deep functions using a total number of parameters similar to the original deep model. We evaluate our method on the TIMIT phoneme recognition task and are able to train shallow fully-connected nets that perform similarly to complex, well-engineered, deep convolutional architectures. Our success in training shallow neural nets to mimic deeper models suggests that there probably exist better algorithms for training shallow feed-forward nets than those currently available.