2016/05/20 by Nikola Milosevic, Nikola Milošević, Milosevic, Nikola +3 · 1 citation
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling #Translation Studies and Practices #cs.AI #cs.CL
paper · pdf · doi:10.48550/arxiv.1605.06319
Phrase modelling, simile extraction, language resource building, crowdsourcing
arxiv created 2016/05/20 · openalex publication_date 2016/05/20 · arxiv updated 2016/05/23 · openalex created_date 2016/06/24 · openalex updated_date 2026/07/28
Similes are natural language expressions used to compare unlikely things, where the comparison is not taken literally. They are often used in everyday communication and are an important part of cultural heritage. Having an up-to-date corpus of similes is challenging, as they are constantly coined and/or adapted to the contemporary times. In this paper we present a methodology for semi-automated collection of similes from the world wide web using text mining techniques. We expanded an existing corpus of traditional similes (containing 333 similes) by collecting 446 additional expressions. We, also, explore how crowdsourcing can be used to extract and curate new similes.