2020/02/25 by Mohammad Jalilpour Monesi, Bernd Accou, Monesi, Mohammad Jalilpour +7 · 3 citations
Engineering · #Audio and Speech Processing (eess.AS) #FOS: Electrical engineering #eess.AS #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2002.10988
3 figures, 6 pages
arxiv created 2020/02/25 · arxiv updated 2020/02/26
Modeling the relationship between natural speech and a recorded electroencephalogram (EEG) helps us understand how the brain processes speech and has various applications in neuroscience and brain-computer interfaces. In this context, so far mainly linear models have been used. However, the decoding performance of the linear model is limited due to the complex and highly non-linear nature of the auditory processing in the human brain. We present a novel Long Short-Term Memory (LSTM)-based architecture as a non-linear model for the classification problem of whether a given pair of (EEG, speech envelope) correspond to each other or not. The model maps short segments of the EEG and the envelope to a common embedding space using a CNN in the EEG path and an LSTM in the speech path. The latter also compensates for the brain response delay. In addition, we use transfer learning to fine-tune the model for each subject. The mean classification accuracy of the proposed model reaches 85%, which is significantly higher than that of a state of the art Convolutional Neural Network (CNN)-based model (73%) and the linear model (69%).