2020/05/18 by Siddique Latif, Muhammad Asim, Latif, Siddique +9 · 1 citation
Computer Science · Psychology · #Audio and Speech Processing (eess.AS) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2005.08447
openalex publication_date 2020/05/18 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Generative adversarial networks (GANs) have shown potential in learning\nemotional attributes and generating new data samples. However, their\nperformance is usually hindered by the unavailability of larger speech emotion\nrecognition (SER) data. In this work, we propose a framework that utilises the\nmixup data augmentation scheme to augment the GAN in feature learning and\ngeneration. To show the effectiveness of the proposed framework, we present\nresults for SER on (i) synthetic feature vectors, (ii) augmentation of the\ntraining data with synthetic features, (iii) encoded features in compressed\nrepresentation. Our results show that the proposed framework can effectively\nlearn compressed emotional representations as well as it can generate synthetic\nsamples that help improve performance in within-corpus and cross-corpus\nevaluation.\n