2020/05/15 by Mostafa M. Mohamed, Björn W. Schuller, Mohamed, Mostafa M. +1 · 1 citation
Computer Science · #Advanced Data Compression Techniques #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Machine Learning (cs.LG) #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2005.07777
openalex publication_date 2020/05/15 · openalex created_date 2022/07/26 · openalex updated_date 2026/07/28
Packet loss is a common problem in data transmission, including speech data\ntransmission. This may affect a wide range of applications that stream audio\ndata, like streaming applications or speech emotion recognition (SER). Packet\nLoss Concealment (PLC) is any technique of facing packet loss. Simple PLC\nbaselines are 0-substitution or linear interpolation. In this paper, we present\na concealment wrapper, which can be used with stacked recurrent neural cells.\nThe concealment cell can provide a recurrent neural network (ConcealNet), that\nperforms real-time step-wise end-to-end PLC at inference time. Additionally,\nextending this with an end-to-end emotion prediction neural network provides a\nnetwork that performs SER from audio with lost frames, end-to-end. The proposed\nmodel is compared against the fore-mentioned baselines. Additionally, a\nbidirectional variant with better performance is utilised. For evaluation, we\nchose the public RECOLA dataset given its long audio tracks with continuous\nemotion labels. ConcealNet is evaluated on the reconstruction of the audio and\nthe quality of corresponding emotions predicted after that. The proposed\nConcealNet model has shown considerable improvement, for both audio\nreconstruction and the corresponding emotion prediction, in environments that\ndo not have losses with long duration, even when the losses occur frequently.\n