2021/04/17 by Laiba Mehnaz, Debanjan Mahata, Mehnaz, Laiba +19
Computer Science · #Artificial Intelligence (cs.AI) #Authorship Attribution and Profiling #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #cs.AI #cs.CL #cs.CY
paper · pdf · doi:10.48550/arxiv.2104.08578
arxiv created 2021/04/17 · openalex publication_date 2021/04/17 · arxiv updated 2021/04/20 · openalex created_date 2021/04/26 · openalex updated_date 2026/07/28
Code-switching is the communication phenomenon where speakers switch between different languages during a conversation. With the widespread adoption of conversational agents and chat platforms, code-switching has become an integral part of written conversations in many multi-lingual communities worldwide. This makes it essential to develop techniques for summarizing and understanding these conversations. Towards this objective, we introduce abstractive summarization of Hindi-English code-switched conversations and develop the first code-switched conversation summarization dataset - GupShup, which contains over 6,831 conversations in Hindi-English and their corresponding human-annotated summaries in English and Hindi-English. We present a detailed account of the entire data collection and annotation processes. We analyze the dataset using various code-switching statistics. We train state-of-the-art abstractive summarization models and report their performances using both automated metrics and human evaluation. Our results show that multi-lingual mBART and multi-view seq2seq models obtain the best performances on the new dataset