vix.ing · top · new · best · stats · spec

Directional MCLP Analysis and Reconstruction for Spatial Speech\n Communication

2021/09/09 by Srikanth Raj Chetupalli, Chetupalli, Srikanth Raj, T.V. Sreenivas +1
Computer Science · Engineering · Neuroscience · #Advanced Adaptive Filtering Techniques #Audio and Speech Processing (eess.AS) #Direction-of-Arrival Estimation Techniques #FOS: Computer and information sciences #FOS: Electrical engineering #Hearing Loss and Rehabilitation #Signal Processing (eess.SP) #Sound (cs.SD) #Speech and Audio Processing #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.2109.04544

openalex publication_date 2021/09/09 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Spatial speech communication, i.e., the reconstruction of spoken signal along\nwith the relative speaker position in the enclosure (reverberation information)\nis considered in this paper. Directional, diffuse components and the source\nposition information are estimated at the transmitter, and perceptually\neffective reproduction is considered at the receiver. We consider spatially\ndistributed microphone arrays for signal acquisition, and node specific signal\nestimation, along with its direction of arrival (DoA) estimation. Short-time\nFourier transform (STFT) domain multi-channel linear prediction (MCLP) approach\nis used to model the diffuse component and relative acoustic transfer function\nis used to model the direct signal component. Distortion-less array response\nconstraint and the time-varying complex Gaussian source model are used in the\njoint estimation of source DoA and the constituent signal components,\nseparately at each node. The intersection between DoA directions at each node\nis used to compute the source position. Signal components computed at the node\nnearest to the estimated source position are taken as the signals for\ntransmission.\n At the receiver, a four channel loud speaker (LS) setup is used for spatial\nreproduction, in which the source spatial image is reproduced relative to a\nchosen virtual listener position in the transmitter enclosure. Vector base\namplitude panning (VBAP) method is used for direct component reproduction using\nthe LS setup and the diffuse component is reproduced equally from all the loud\nspeakers after decorrelation. This scheme of spatial speech communication is\nshown to be effective and more natural for hands-free telecommunication,\nthrough either loudspeaker listening or binaural headphone listening with head\nrelated transfer function (HRTF) based presentation.\n

Citations

Related