2016/10/03 by Anoop Kunchukuttan, Pushpak Bhattacharyya, Kunchukuttan, Anoop +1
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Handwritten Text Recognition Techniques #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1610.00634
openalex publication_date 2016/10/03 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We explore the use of the orthographic syllable, a variable-length consonant-vowel sequence, as a basic unit of translation between related languages which use abugida or alphabetic scripts. We show that orthographic syllable level translation significantly outperforms models trained over other basic units (word, morpheme and character) when training over small parallel corpora.