vix.ing · top · new · best · stats · spec

HUMBO: Bridging Response Generation and Facial Expression Synthesis

2019/05/24 by Shang‐Yu Su, Su, Shang-Yu, Po‐Wei Lin +3
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition #Multimedia (cs.MM) #Multimodal Machine Learning Applications

paper · pdf · doi:10.48550/arxiv.1905.11240

openalex publication_date 2019/05/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Spoken dialogue systems that assist users to solve complex tasks such as movie ticket booking have become an emerging research topic in artificial intelligence and natural language processing areas. With a well-designed dialogue system as an intelligent personal assistant, people can accomplish certain tasks more easily via natural language interactions. Today there are several virtual intelligent assistants in the market; however, most systems only focus on textual or vocal interaction. In this paper, we present HUMBO, a system aiming at generating dialogue responses and simultaneously synthesize corresponding visual expressions on faces for better multimodal interaction. HUMBO can (1) let users determine the appearances of virtual assistants by a single image, and (2) generate coherent emotional utterances and facial expressions on the user-provided image. This is not only a brand new research direction but more importantly, an ultimate step toward more human-like virtual assistants.

Citations

Related