2022/12/07 by Shih-Hong Huang, Chieh-Yang Huang, Huang, Shih-Hong +9
Computer Science · #Context-Aware Activity Recognition Systems #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Mobile Crowdsensing and Crowdsourcing #Speech and dialogue systems
paper · pdf · doi:10.48550/arxiv.2212.03969
openalex publication_date 2022/12/07 · openalex created_date 2022/12/22 · openalex updated_date 2026/07/28
Real-time crowd-powered systems, such as Chorus/Evorus, VizWiz, and Apparition, have shown how incorporating humans into automated systems could supplement where the automatic solutions fall short. However, one unspoken bottleneck of applying such architectures to more scenarios is the longer latency of including humans in the loop of automated systems. For the applications that have hard constraints in turnaround times, human-operated components' longer latency and large speed variation seem to be apparent deal breakers. This paper explicates and quantifies these limitations by using a human-powered text-based backend to hold conversations with users through a voice-only smart speaker. Smart speakers must respond to users' requests within seconds, so the workers behind the scenes only have a few seconds to compose answers. We measured the end-to-end system latency and the conversation quality with eight pairs of participants, showing the challenges and superiority of such systems.