2022/10/02 by Gavin Abercrombie, Verena Rieser, Abercrombie, Gavin +1 · 1 citation
Computer Science · Medicine · #AI in Service Interactions #Artificial Intelligence in Healthcare and Education #Computation and Language (cs.CL) #FOS: Computer and information sciences #Topic Modeling #cs.CL
paper · pdf · doi:10.48550/arxiv.2210.00572
Accepted for publication at AACL 2022
arxiv created 2022/10/02 · openalex publication_date 2022/10/02 · arxiv updated 2022/10/04 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Conversational AI systems can engage in unsafe behaviour when handling users' medical queries that can have severe consequences and could even lead to deaths. Systems therefore need to be capable of both recognising the seriousness of medical inputs and producing responses with appropriate levels of risk. We create a corpus of human written English language medical queries and the responses of different types of systems. We label these with both crowdsourced and expert annotations. While individual crowdworkers may be unreliable at grading the seriousness of the prompts, their aggregated labels tend to agree with professional opinion to a greater extent on identifying the medical queries and recognising the risk types posed by the responses. Results of classification experiments suggest that, while these tasks can be automated, caution should be exercised, as errors can potentially be very serious.