2021/01/27 by Premjeet Singh, Goutam Saha, Sahidullah · 1 citation
Computer Science · Psychology · Mathematics · #Speech and Audio Processing #Music and Audio Processing #Emotion and Mood Recognition #Short-time Fourier transform #Computer science #Speech recognition #Pattern recognition (psychology) #Artificial intelligence #Time–frequency analysis #Spectrogram #Fourier transform #Image warping #Classifier (UML) #Transformation (genetics) #Frequency domain #Mathematics #Fourier analysis #Computer vision
paper · doi:10.1109/iccci50826.2021.9402569
openalex publication_date 2021/01/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
In this work, we explore the constant-Q transform (CQT) for speech emotion recognition (SER). The CQT-based time-frequency analysis provides variable spectro-temporal resolution with higher frequency resolution at lower frequencies. Since lower-frequency regions of speech signal contain more emotion-related information than higher-frequency regions, the increased low-frequency resolution of CQT makes it more promising for SER than standard short-time Fourier transform (STFT). We present a comparative analysis of short-term acoustic features based on STFT and CQT for SER with deep neural network (DNN) as a back-end classifier. We optimize different parameters for both features. The CQT-based features outperform the STFT-based spectral features for SER experiments. Further experiments with cross-corpora evaluation demonstrate that the CQT-based systems provide better generalization with out-of-domain training data.