vix.ing · top · new · best · stats · spec

Creating a contemporary corpus of similes in Serbian by using natural language processing

2018/11/22 by Nikola Milosevic, Nikola Milošević, Milosevic, Nikola +3
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Lexicography and Language Studies #Machine Learning (cs.LG) #Natural Language Processing Techniques #cs.AI #cs.CL #cs.CY #cs.LG #linguistics and terminology studies

paper · pdf · doi:10.48550/arxiv.1811.10422

15 pages, submitted to journal Slovo, however, later withdrawn to correct. Additional work was not done on it, so it is still waiting to be extended. Output of the system can be seen here: http://ezbirka.starisloveni.com/. arXiv admin note: text overlap with arXiv:1605.06319

arxiv created 2018/11/22 · openalex publication_date 2018/11/22 · arxiv updated 2018/11/27 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

Simile is a figure of speech that compares two things through the use of connection words, but where comparison is not intended to be taken literally. They are often used in everyday communication, but they are also a part of linguistic cultural heritage. In this paper we present a methodology for semi-automated collection of similes from the World Wide Web using text mining and machine learning techniques. We expanded an existing corpus by collecting 442 similes from the internet and adding them to the existing corpus collected by Vuk Stefanovic Karadzic that contained 333 similes. We, also, introduce crowdsourcing to the collection of figures of speech, which helped us to build corpus containing 787 unique similes.

Citations

Related