2025/06/07 by Vahid Dastjerdi, Hossein
#AI chatbots #Conversational AI #Multimodal interaction #Text #User engagemen #User satisfaction #Visual inputs #Voice
paper · doi:10.57647/jntell.2025.0401.05
This holistic study took a critical look at the integration of various modes of communication, particularly text, voice, and visual inputs, for the development of a seamless and cohesive multimodal conversational experience within AI chatbots. The focus of this research was directed toward Iranian EFL intermediate high school students aged between 15 and 19 years. It is important to note that the existing conversational AI systems usually rely on a single mode of interaction, which in effect seriously limits their overall effectiveness and usability. By integrating text, voice, and visual elements in this innovative approach, this research aims at increasing user engagement and satisfaction levels among the students participating in the study. A sample of 200 male and female students was conveniently selected and engaged in multiple interactions with custom-developed AI chatbots over a period of three months. Each subject experienced text-only, voice-only, visual-only, and multimodal interactions in a random order. Data collections included interaction duration, frequency, and satisfaction surveys, while further data was collected through focus groups. Quantitative analysis using MANOVA and qualitative thematic analysis have together shed some very important light on multimodal interaction. These interactions, having been shown to significantly heighten the overall user experience, promise new directions in the further development of conversational AI technologies.