2016/04/28 by Yin Xian, Xian, Yin, Andrew Thompson +9
Computer Science · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Music and Audio Processing #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #cs.LG #cs.SD
paper · pdf · doi:10.48550/arxiv.1605.01755
22 figures
arxiv created 2016/04/28 · openalex publication_date 2016/04/28 · arxiv updated 2016/05/09 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We introduce the use of DCTNet, an efficient approximation and alternative to PCANet, for acoustic signal classification. In PCANet, the eigenfunctions of the local sample covariance matrix (PCA) are used as filterbanks for convolution and feature extraction. When the eigenfunctions are well approximated by the Discrete Cosine Transform (DCT) functions, each layer of of PCANet and DCTNet is essentially a time-frequency representation. We relate DCTNet to spectral feature representation methods, such as the the short time Fourier transform (STFT), spectrogram and linear frequency spectral coefficients (LFSC). Experimental results on whale vocalization data show that DCTNet improves classification rate, demonstrating DCTNet's applicability to signal processing problems such as underwater acoustics.