Comparison of parametric representations for monosyllabic word recognition in continuously spoken sentences
1980/08/01 by S. Davis, P. Mermelstein · 5,365 citations
Computer Science · Mathematics · #Artificial intelligence #Cepstrum #Computer science #Dynamic time warping #Feature extraction #Linear prediction #Linguistics #Mathematics #Mel-frequency cepstrum #Music and Audio Processing #Parametric statistics #Pattern recognition (psychology) #Set (abstract data type) #Speech Recognition and Synthesis #Speech recognition #Statistics #Syllable #Time Series Analysis and Forecasting #Word (group theory) #Word recognition
paper · doi:10.1109/tassp.1980.1163420
published in IEEE Transactions on Acoustics Speech and Signal Processing 28(4), 357-366 (Institute of Electrical and Electronics Engineers)
openalex publication_date 1980/08/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/02
Abstract
Several parametric representations of the acoustic signal were compared with regard to word recognition performance in a syllable-oriented continuous speech recognition system. The vocabulary included many phonetically similar monosyllabic words, therefore the emphasis was on the ability to retain phonetically significant acoustic information in the face of syntactic and duration variations. For each parameter set (based on a mel-frequency cepstrum, a linear frequency cepstrum, a linear prediction cepstrum, a linear prediction spectrum, or a set of reflection coefficients), word templates were generated using an efficient dynamic warping method, and test data were time registered with the templates. A set of ten mel-frequency cepstrum coefficients computed every 6.4 ms resulted in the best performance, namely 96.5 percent and 95.0 percent recognition with each of two speakers. The superior performance of the mel-frequency cepstrum coefficients may be attributed to the fact that they better represent the perceptually relevant aspects of the short-term speech spectrum.
Citations
Cited by
- Incorporation of Speech Duration Information in Score Fusion of Speaker Recognition Systems
- Learning from Between-class Examples for Deep Sound Recognition
- Deep generative variational autoencoding for replay spoof detection in automatic speaker verification
- Active Mini-Batch Sampling using Repulsive Point Processes
- LEAF: A Learnable Frontend for Audio Classification
- Pre-training in Deep Reinforcement Learning for Automatic Speech Recognition
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- On the Exploitability of Audio Machine Learning Pipelines to Surreptitious Adversarial Examples
- A neurocomputational account of taxonomic responding and fast mapping in early word learning.
- Using NLP to analyze whether customer statements comply with their inner belief
- Robust Support Vector Machines for Speaker Verification Task
- Emotional speech recognition: Resources, features, and methods
- Recognising realistic emotions and affect in speech: State of the art and lessons learnt from the first challenge
- Mean Hilbert envelope coefficients (MHEC) for robust speaker and language identification
- An overview of text-independent speaker recognition: From features to supervectors
- Viterbi Extraction tutorial with Hidden Markov Toolkit
- An Empirical Study on End-to-End Singing Voice Synthesis with Encoder-Decoder Architectures
- The Minor Fall, the Major Lift: Inferring Emotional Valence of Musical Chords through Lyrics
- Combining evidences from magnitude and phase information using VTEO for person recognition using humming
- Interpretable Filter Learning Using Soft Self-attention For Raw Waveform Speech Recognition
- End-to-end Phoneme Sequence Recognition using Convolutional Neural Networks
- Generating Video Descriptions with Topic Guidance
- Audio-Visual Self-Supervised Terrain Type Discovery for Mobile Platforms
- Adversarial Example Detection by Classification for Deep Speech Recognition
- Modelling of Musical Perception using Spectral Knowledge Representation
- Choice of Mel Filter Bank in Computing MFCC of a Resampled Speech
- Deep Spiking Neural Networks for Large Vocabulary Automatic Speech Recognition
- African elephants address one another with individually specific name-like calls
- Fidelity of fricative measurements in remote data collection
- A Unified Deep Speaker Embedding Framework for Mixed-Bandwidth Speech Data
- Music Genre Classification using Machine Learning Techniques
- Practical Selection of SVM Supervised Parameters with Different Feature Representations for Vowel Recognition
- On combining features for single-channel robust speech recognition in reverberant environments
- Discrimination between mothers’ infant- and adult-directed speech using hidden Markov models
- The importance of phase in speech enhancement
- An Overview of Lead and Accompaniment Separation in Music
- End-to-end acoustic modeling using convolutional neural networks for HMM-based automatic speech recognition
- Learnable MFCCs for Speaker Verification
- Classification of Audio Segments in Call Center Recordings using Convolutional Recurrent Neural Networks
- Consensus-based Sequence Training for Video Captioning
- Speech Emotion Recognition Based on Multi-feature and Multi-lingual Fusion
- Feature-level and Model-level Audiovisual Fusion for Emotion Recognition in the Wild
- Locality-Sensitive Hashing with Margin Based Feature Selection
- A Subband-Based SVM Front-End for Robust ASR
- Keyword Mamba: Spoken Keyword Spotting with State Space Models
- nnAudio: An on-the-fly GPU Audio to Spectrogram Conversion Toolbox Using 1D Convolution Neural Networks
- Deception Detection in Videos
- DeepGB-TB: A Risk-Balanced Cross-Attention Gradient-Boosted Convolutional Network for Rapid, Interpretable Tuberculosis Screening
- Training for Speech Recognition on Coprocessors
- Automatic Long-Term Deception Detection in Group Interaction Videos
- Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
- cvpaper.challenge in 2016: Futuristic Computer Vision through 1,600 Papers Survey
- Deep Generative Variational Autoencoding for Replay Spoof Detection in Automatic Speaker Verification
- Feature Trajectory Dynamic Time Warping for Clustering of Speech Segments
- Abnormal noise monitoring of subway vehicles based on combined acoustic features
- Deep Reinforcement Learning with Pre-training for Time-efficient Training of Automatic Speech Recognition
- Spoken language identification: An overview of past and present research trends
- VisemeNet: Audio-Driven Animator-Centric Speech Animation
- Identification of fake stereo audio
- Audio Content Analysis
- LISTEN: Lightweight Industrial Sound-representable Transformer for Edge Notification
- Multimodal Speech Emotion Recognition and Ambiguity Resolution
- Unsupervised Spoken Term Discovery on Untranscribed Speech
- Modified Mel Filter Bank to Compute MFCC of Subsampled Speech
- Emotion Invariant Speaker Embeddings for Speaker Identification with Emotional Speech
- HLT-NUS Submission for NIST 2019 Multimedia Speaker Recognition Evaluation
- Hyperplane Arrangements and Locality-Sensitive Hashing with Lift
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions
- WSNet: Compact and Efficient Networks Through Weight Sampling
- Improving Visual Recognition using Ambient Sound for Supervision
- Advances in Small-Footprint Keyword Spotting: A Comprehensive Review of Efficient Models and Algorithms
- Raw Waveform Encoder with Multi-Scale Globally Attentive Locally Recurrent Networks for End-to-End Speech Recognition
- Unsupervised heart abnormality detection based on phonocardiogram analysis with Beta Variational Auto-Encoders
- An Empirical Comparison of SVM and Some Supervised Learning Algorithms for Vowel recognition
- An Information Retrieval Approach to Finding Dependent Subspaces of Multiple Views
- Sensor Transformation Attention Networks
- Experiments on Open-Set Speaker Identification with Discriminatively Trained Neural Networks
- On desensitizing the Mel-Cepstrum to spurious spectral components for Robust Speech Recognition
- DeepDrummer : Generating Drum Loops using Deep Learning and a Human in the Loop
- Recognizing Abnormal Heart Sounds Using Deep Learning
- Speech recognition with quaternion neural networks
- Data-driven audio recognition: a supervised dictionary approach
- Towards Learning to Speak and Hear Through Multi-Agent Communication\n over a Continuous Acoustic Channel
- A machine learning perspective on the emotional content of Parkinsonian speech
- Comparison of different implementations of MFCC
- Emotions in the Loop: A Survey of Affective Computing for Emotional Support
- Parallel training of DNNs with Natural Gradient and Parameter Averaging
- Cepstral noise subtraction for robust automatic speech recognition
- A snippet in a snippet: Development of the Matryoshka principle for the construction of very short musical stimuli (plinks)
- A comparison of speaker identification results using features based on cepstrum and Fourier-Bessel expansion
- Using Deep Learning for Detecting Spoofing Attacks on Speech Signals
- PEFAC - A Pitch Estimation Algorithm Robust to High Levels of Noise
- Lung sound classification using cepstral-based statistical features
- Local spectral variability features for speaker verification
- Analysis of the correlation structure for a neural predictive model with application to speech recognition
- Developments and directions in speech recognition and understanding, Part 1 [DSP Education]
- Emotion recognition from Assamese speeches using MFCC features and GMM classifier
- Probabilistic Linear Discriminant Analysis for Acoustic Modeling
- Fractal dimensions of speech sounds: Computation and application to automatic speech recognition
- Blind normalization of speech from different channels
- Robust text-independent speaker identification using Gaussian mixture speaker models
- Semi-supervised Phoneme Recognition with Recurrent Ladder Networks
- Talking condition recognition in stressful and emotional talking environments based on CSPHMM2s
- Effects of Number of Filters of Convolutional Layers on Speech Recognition Model Accuracy
- Interpretable Representation Learning for Speech and Audio Signals Based on Relevance Weighting
- Multimodal Transfer Deep Learning with Applications in Audio-Visual Recognition
- Affective computing using speech and eye gaze: a review and bimodal system proposal for continuous affect prediction
- Environmental sound classification using convolution neural networks with different integrated loss functions
- Novel dual-channel long short-term memory compressed capsule networks for emotion recognition
- Constant Q cepstral coefficients: A spoofing countermeasure for automatic speaker verification
Related