vix.ing · top · new · best · stats · spec

Voice Activity Projection Model with Multimodal Encoders

2025/06/04 by Takeshi Saga, Saga, Takeshi, Catherine Pélachaud +1
Computer Science · Psychology · #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #Phonetics and Phonology Research #Speech and Audio Processing

paper · pdf · doi:10.48550/arxiv.2506.03980

openalex publication_date 2025/06/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Turn-taking management is crucial for any social interaction. Still, it is challenging to model human-machine interaction due to the complexity of the social context and its multimodal nature. Unlike conventional systems based on silence duration, previous existing voice activity projection (VAP) models successfully utilized a unified representation of turn-taking behaviors as prediction targets, which improved turn-taking prediction performance. Recently, a multimodal VAP model outperformed the previous state-of-the-art model by a significant margin. In this paper, we propose a multimodal model enhanced with pre-trained audio and face encoders to improve performance by capturing subtle expressions. Our model performed competitively, and in some cases, even better than state-of-the-art models on turn-taking metrics. All the source codes and pretrained models are available at https://github.com/sagatake/VAPwithAudioFaceEncoders.

Related