2017/05/13 by Jinfeng Rao, Ferhan Ture, Rao, Jinfeng +8
Computer Science · #FOS: Computer and information sciences #Information Retrieval (cs.IR) #Music and Audio Processing #Speech Recognition and Synthesis #Topic Modeling #cs.IR
paper · pdf · doi:10.48550/arxiv.1705.04892
arxiv created 2017/05/13 · openalex publication_date 2017/05/13 · arxiv updated 2017/05/16 · openalex created_date 2019/06/27 · openalex updated_date 2026/07/28
We tackle the novel problem of navigational voice queries posed against an entertainment system, where viewers interact with a voice-enabled remote controller to specify the program to watch. This is a difficult problem for several reasons: such queries are short, even shorter than comparable voice queries in other domains, which offers fewer opportunities for deciphering user intent. Furthermore, ambiguity is exacerbated by underlying speech recognition errors. We address these challenges by integrating word- and character-level representations of the queries and by modeling voice search sessions to capture the contextual dependencies in query sequences. Both are accomplished with a probabilistic framework in which recurrent and feedforward neural network modules are organized in a hierarchical manner. From a raw dataset of 32M voice queries from 2.5M viewers on the Comcast Xfinity X1 entertainment system, we extracted data to train and test our models. We demonstrate the benefits of our hybrid representation and context-aware model, which significantly outperforms models without context as well as the current deployed product.