2021/11/24 by Adam Farris, Farris, Adam, Aryaman Arora +1
Arts and Humanities · Computer Science · Social Sciences · #68T50 #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Linguistic Variation and Morphology #Natural Language Processing Techniques #Syntax, Semantics, Linguistic Variation
paper · pdf · doi:10.48550/arxiv.2111.12783
openalex publication_date 2021/11/24 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present the first linguistically annotated treebank of Ashokan Prakrit, an early Middle Indo-Aryan dialect continuum attested through Emperor Ashoka Maurya's 3rd century BCE rock and pillar edicts. For annotation, we used the multilingual Universal Dependencies (UD) formalism, following recent UD work on Sanskrit and other Indo-Aryan languages. We touch on some interesting linguistic features that posed issues in annotation: regnal names and other nominal compounds, "proto-ergative" participial constructions, and possible grammaticalizations evidenced by sandhi (phonological assimilation across morpheme boundaries). Eventually, we plan for a complete annotation of all attested Ashokan texts, towards the larger goals of improving UD coverage of different diachronic stages of Indo-Aryan and studying language change in Indo-Aryan using computational methods.