2017/05/25 by Amirsina Torfi, Jeremy Dawson, Torfi, Amirsina +3
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Speech and Audio Processing
paper · pdf · doi:10.48550/arxiv.1705.09422
openalex publication_date 2017/05/25 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN)\narchitecture has been proposed for speaker verification in the text-independent\nsetting. One of the main challenges is the creation of the speaker models. Most\nof the previously-reported approaches create speaker models based on averaging\nthe extracted features from utterances of the speaker, which is known as the\nd-vector system. In our paper, we propose an adaptive feature learning by\nutilizing the 3D-CNNs for direct speaker model creation in which, for both\ndevelopment and enrollment phases, an identical number of spoken utterances per\nspeaker is fed to the network for representing the speakers' utterances and\ncreation of the speaker model. This leads to simultaneously capturing the\nspeaker-related information and building a more robust system to cope with\nwithin-speaker variation. We demonstrate that the proposed method significantly\noutperforms the traditional d-vector verification system. Moreover, the\nproposed system can also be an alternative to the traditional d-vector system\nwhich is a one-shot speaker modeling system by utilizing 3D-CNNs.\n