vix.ing · top · new · best · stats · spec

Vulgaris: Analysis of a Corpus for Middle-Age Varieties of Italian\n Language

2020/10/12 by Andrea Zugarini, Zugarini, Andrea, Matteo Tiezzi +3
Computer Science · Social Sciences · #Artificial Intelligence (cs.AI) #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Linguistic Variation and Morphology #Natural Language Processing Techniques

paper · pdf · doi:10.48550/arxiv.2010.05993

openalex publication_date 2020/10/12 · openalex created_date 2022/07/25 · openalex updated_date 2026/07/28

Abstract

Italian is a Romance language that has its roots in Vulgar Latin. The birth\nof the modern Italian started in Tuscany around the 14th century, and it is\nmainly attributed to the works of Dante Alighieri, Francesco Petrarca and\nGiovanni Boccaccio, who are among the most acclaimed authors of the medieval\nage in Tuscany. However, Italy has been characterized by a high variety of\ndialects, which are often loosely related to each other, due to the past\nfragmentation of the territory. Italian has absorbed influences from many of\nthese dialects, as also from other languages due to dominion of portions of the\ncountry by other nations, such as Spain and France. In this work we present\nVulgaris, a project aimed at studying a corpus of Italian textual resources\nfrom authors of different regions, ranging in a time period between 1200 and\n1600. Each composition is associated to its author, and authors are also\ngrouped in families, i.e. sharing similar stylistic/chronological\ncharacteristics. Hence, the dataset is not only a valuable resource for\nstudying the diachronic evolution of Italian and the differences between its\ndialects, but it is also useful to investigate stylistic aspects between single\nauthors. We provide a detailed statistical analysis of the data, and a\ncorpus-driven study in dialectology and diachronic varieties.\n

Citations

Related