2019/12/11 by Masaki Oguni, Oguni, Masaki, Yohei Seki +3
Computer Science · Mathematics · #Artificial intelligence #Character (mathematics) #Computer science #Embedding #FOS: Computer and information sciences #Geography #Information Retrieval (cs.IR) #Information retrieval #Language model #Linguistics #Mathematics #Natural Language Processing Techniques #Natural language processing #Recipe #Software Engineering Research #Topic Modeling #Word (group theory) #Word embedding #cs.IR #n-gram
paper · pdf · doi:10.48550/arxiv.1912.05171
published in arXiv (Cornell University) (Cornell University) · 5 pages, 2 figures
openalex publication_date 2019/12/11 · arxiv created 2019/12/21 · arxiv updated 2019/12/24 · openalex created_date 2019/12/26 · openalex updated_date 2026/07/28
In user-generated recipe websites, users post their-original recipes. Some recipes, however, are very similar in major components such as the cooking instructions to other recipes. We refer to such recipes as "near-duplicate recipes". In this study, we propose a method that extends the "Word Mover's Distance", which calculates distances between texts based on word embedding, to character 3-gram embedding. Using a corpus of over 1.21 million recipes, we learned the word embedding and the character 3-gram embedding by using a Skip-Gram model with negative sampling and fastText to extract candidate pairs of near-duplicate recipes. We then annotated these candidates and evaluated the proposed method against a comparison method. Our results demonstrated that near-duplicate recipes that were not detected by the comparison method were successfully detected by the proposed method.