2017/05/29 by Michał Zapotoczny, Zapotoczny, Michał, Paweł Rychlikowski +3
Computer Science · #Artificial intelligence #Artificial neural network #Computation and Language (cs.CL) #Computer science #Dependency (UML) #Dependency grammar #FOS: Computer and information sciences #Linguistics #Machine Learning (cs.LG) #Natural Language Processing Techniques #Natural language processing #Neural and Evolutionary Computing (cs.NE) #Parsing #Text Readability and Simplification #Topic Modeling #Word (group theory) #cs.CL #cs.LG #cs.NE
paper · pdf · doi:10.48550/arxiv.1705.10209
published in arXiv (Cornell University) (Cornell University) · preprint accepted into the TSD2017
arxiv created 2017/05/29 · openalex publication_date 2017/05/29 · arxiv updated 2017/05/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
We show that a recently proposed neural dependency parser can be improved by joint training on multiple languages from the same family. The parser is implemented as a deep neural network whose only input is orthographic representations of words. In order to successfully parse, the network has to discover how linguistically relevant concepts can be inferred from word spellings. We analyze the representations of characters and words that are learned by the network to establish which properties of languages were accounted for. In particular we show that the parser has approximately learned to associate Latin characters with their Cyrillic counterparts and that it can group Polish and Russian words that have a similar grammatical function. Finally, we evaluate the parser on selected languages from the Universal Dependencies dataset and show that it is competitive with other recently proposed state-of-the art methods, while having a simple structure.