2018/01/01 by Amrith Krishna, Krishna, Amrith, Bishal Santra +12 · 1 citation
Computer Science · Mathematics · #Artificial intelligence #Computation and Language (cs.CL) #Computer science #Context (archaeology) #FOS: Computer and information sciences #Graph #Linguistics #Mathematics #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Natural language processing #Parsing #Sanskrit #Segmentation #Sentence #Speech recognition #Task (project management) #Text Readability and Simplification #Text segmentation #Theoretical computer science #Topic Modeling #Word (group theory) #cs.CL
paper · pdf · doi:10.48550/arxiv.1809.01446
published in arXiv (Cornell University) (Cornell University) · version 2: Corrected typo in Table1, page7 | Accepted in EMNLP 2018. Supplementary material can be found at - http://cse.iitkgp.ac.in/~amrithk/1080_supp.pdf
openalex publication_date 2018/09/05 · arxiv created 2018/10/25 · arxiv updated 2018/10/26 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
The configurational information in sentences of a free word order language\nsuch as Sanskrit is of limited use. Thus, the context of the entire sentence\nwill be desirable even for basic processing tasks such as word segmentation. We\npropose a structured prediction framework that jointly solves the word\nsegmentation and morphological tagging tasks in Sanskrit. We build an energy\nbased model where we adopt approaches generally employed in graph based parsing\ntechniques (McDonald et al., 2005a; Carreras, 2007). Our model outperforms the\nstate of the art with an F-Score of 96.92 (percentage improvement of 7.06%)\nwhile using less than one-tenth of the task-specific training data. We find\nthat the use of a graph based ap- proach instead of a traditional lattice-based\nsequential labelling approach leads to a percentage gain of 12.6% in F-Score\nfor the segmentation task.\n