Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
1980/08/01 by S. Davis, P. Mermelstein · 60 citations
Computer Science · #Music and Audio Processing #Speech Recognition and Synthesis #Time Series Analysis and Forecasting
paper · doi:10.1109/tassp.1980.1163420
Abstract
Several parametric representations of the acoustic signal were compared with regard to word recognition performance in a syllable-oriented continuous speech recognition system. The vocabulary included many phonetically similar monosyllabic words, therefore the emphasis was on the ability to retain phonetically significant acoustic information in the face of syntactic and duration variations. For each parameter set (based on a mel-frequency cepstrum, a linear frequency cepstrum, a linear prediction cepstrum, a linear prediction spectrum, or a set of reflection coefficients), word templates were generated using an efficient dynamic warping method, and test data were time registered with the templates. A set of ten mel-frequency cepstrum coefficients computed every 6.4 ms resulted in the best performance, namely 96.5 percent and 95.0 percent recognition with each of two speakers. The superior performance of the mel-frequency cepstrum coefficients may be attributed to the fact that they better represent the perceptually relevant aspects of the short-term speech spectrum.
Citations
Cited by
- Incorporation of Speech Duration Information in Score Fusion of Speaker\n Recognition Systems
- Learning from Between-class Examples for Deep Sound Recognition
- Deep generative variational autoencoding for replay spoof detection in automatic speaker verification
- Active Mini-Batch Sampling using Repulsive Point Processes
- LEAF: A Learnable Frontend for Audio Classification
- Pre-training in Deep Reinforcement Learning for Automatic Speech Recognition
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- On the Exploitability of Audio Machine Learning Pipelines to Surreptitious Adversarial Examples
- A neurocomputational account of taxonomic responding and fast mapping in early word learning.
- Using NLP to analyze whether customer statements comply with their inner\n belief
- Robust Support Vector Machines for Speaker Verification Task
- Emotional speech recognition: Resources, features, and methods
- Recognising realistic emotions and affect in speech: State of the art and lessons learnt from the first challenge
- Mean Hilbert envelope coefficients (MHEC) for robust speaker and language identification
- An overview of text-independent speaker recognition: From features to supervectors
- Viterbi Extraction tutorial with Hidden Markov Toolkit
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- Pretoriana, no. 035, April 1961
- Combining evidences from magnitude and phase information using VTEO for person recognition using humming
- Interpretable Filter Learning Using Soft Self-attention For Raw Waveform Speech Recognition
- End-to-end Phoneme Sequence Recognition using Convolutional Neural\n Networks
- Generating Video Descriptions with Topic Guidance
- Audio-Visual Self-Supervised Terrain Type Discovery for Mobile Platforms
- Adversarial Example Detection by Classification for Deep Speech Recognition
- Modelling of Musical Perception using Spectral Knowledge Representation
- Choice of Mel Filter Bank in Computing MFCC of a Resampled Speech
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- African elephants address one another with individually specific name-like calls
- Fidelity of fricative measurements in remote data collection
- A Unified Deep Speaker Embedding Framework for Mixed-Bandwidth Speech Data
- Music Genre Classification using Machine Learning Techniques
- Practical Selection of SVM Supervised Parameters with Different Feature Representations for Vowel Recognition
- On combining features for single-channel robust speech recognition in reverberant environments
- Discrimination between mothers’ infant- and adult-directed speech using hidden Markov models
- The importance of phase in speech enhancement
- An Overview of Lead and Accompaniment Separation in Music
- End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition
- Learnable MFCCs for Speaker Verification
- Classification of Audio Segments in Call Center Recordings using\n Convolutional Recurrent Neural Networks
- Consensus-based Sequence Training for Video Captioning
- Speech Emotion Recognition Based on Multi-feature and Multi-lingual Fusion
- Feature-level and Model-level Audiovisual Fusion for Emotion Recognition in the Wild
- Locality-Sensitive Hashing with Margin Based Feature Selection
- A Subband-Based SVM Front-End for Robust ASR
- Keyword Mamba: Spoken Keyword Spotting with State Space Models
- nnAudio: An on-the-fly GPU Audio to Spectrogram Conversion Toolbox Using 1D Convolution Neural Networks
- Deception Detection in Videos
- DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening
- Training for Speech Recognition on Coprocessors
- Automatic Long-Term Deception Detection in Group Interaction Videos
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Deep Generative Variational Autoencoding for Replay Spoof Detection in\n Automatic Speaker Verification
- Feature Trajectory Dynamic Time Warping for Clustering of Speech Segments
- Abnormal noise monitoring of subway vehicles based on combined acoustic features
- Deep Reinforcement Learning with Pre-training for Time-efficient Training of Automatic Speech Recognition
- Spoken language identification: An overview of past and present research trends
- VisemeNet: Audio-Driven Animator-Centric Speech Animation
- Identification of fake stereo audio
- Audio Content Analysis
- LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
Related