2003/04/10 by David Levin, David N. Levin
Computer Science · Mathematics · #Arithmetic #Artificial intelligence #Blind Source Separation Techniques #Cepstrum #Channel (broadcasting) #Computer science #Feature extraction #Impulse response #Invertible matrix #LTI system theory #Linear prediction #Linear system #Mathematical analysis #Mathematics #Mel-frequency cepstrum #Music and Audio Processing #Normalization (sociology) #Speech and Audio Processing #Speech recognition #Subtraction #Telecommunications #cs.CL
paper · pdf · doi:10.1121/1.1755235
25 pages, 7 figures
arxiv created 2003/04/10 · openalex publication_date 2004/07/01 · arxiv updated 2009/11/30 · openalex created_date 2017/10/20 · openalex updated_date 2026/08/05
We show how to construct a channel-independent representation of speech that has propagated through a noisy reverberant channel. This is done by blindly rescaling the cepstral time series by a nonlinear function, with the form of this scale function being determined by previously encountered cepstra from that channel. The rescaled form of the time series is an invariant property of it in the following sense: It is unaffected if the time series is transformed by any time-independent invertible distortion. Because a linear channel with stationary noise and impulse response transforms cepstra in this way, the new technique can be used to remove the channel dependence of a cepstral time series. In experiments, the method achieved greater channel-independence than cepstral mean normalization, and it was comparable to the combination of cepstral mean normalization and spectral subtraction, despite the fact that no measurements of channel noise or reverberations were required (unlike spectral subtraction).