2020/10/23 by Arij Riabi, Thomas Scialom, Riabi, Arij +9 · 27 citations
Computer Science · #Algorithm #Artificial intelligence #Computer science #Data collection #Deep learning #Information retrieval #Labeled data #Language model #Linguistics #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #One shot #Question answering #Shot (pellet) #State (computer science) #Task (project management) #Topic Modeling #Training set #Zero (linguistics) #cs.CL
paper · pdf · doi:10.48550/arxiv.2010.12643
published in arXiv (Cornell University) (Cornell University) · 7 pages
openalex publication_date 2020/10/23 · arxiv created 2021/10/14 · arxiv updated 2021/10/15 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/08
Coupled with the availability of large scale datasets, deep learning\narchitectures have enabled rapid progress on the Question Answering task.\nHowever, most of those datasets are in English, and the performances of\nstate-of-the-art multilingual models are significantly lower when evaluated on\nnon-English data. Due to high data collection costs, it is not realistic to\nobtain annotated data for each language one desires to support.\n We propose a method to improve the Cross-lingual Question Answering\nperformance without requiring additional annotated data, leveraging Question\nGeneration models to produce synthetic samples in a cross-lingual fashion. We\nshow that the proposed method allows to significantly outperform the baselines\ntrained on English data only. We report a new state-of-the-art on four\nmultilingual datasets: MLQA, XQuAD, SQuAD-it and PIAF (fr).\n