2006/09/17 by Daniel Neiberg, Kjell Elenius, Kornel Laskowski · 2 citations
Computer Science · Psychology · #Speech Recognition and Synthesis #Speech and Audio Processing #Emotion and Mood Recognition #Computer science #Speech recognition #Emotion recognition #Speaker recognition #Mixture model #Artificial intelligence #Natural language processing
paper · doi:10.21437/interspeech.2006-277
openalex publication_date 2006/09/17 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Automatic detection of emotions has been evaluated using standard Mel-frequency Cepstral Coefficients, MFCCs, and a variant, MFCC-low, calculated between 20 and 300 Hz, in order to model pitch. Also plain pitch features have been used. These acoustic features have all been modeled by Gaussian mixture models, GMMs, on the frame level. The method has been tested on two different corpora and languages; Swedish voice controlled telephone services and English meetings. The results indicate that using GMMs on the frame level is a feasible technique for emotion classification. The two MFCC methods have similar performance, and MFCC-low outperforms the pitch features. Combining the three classifiers significantly improves performance. 1.