vix.ing · top · new · best · stats · spec

Neural Semi-Markov Conditional Random Fields for Robust Character-Based\n Part-of-Speech Tagging

2018/08/13 by Apostolos Kemos, Kemos, Apostolos, Heike Adel +3 · 1 citation
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech Recognition and Synthesis #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1808.04208

openalex publication_date 2018/08/13 · openalex created_date 2022/08/04 · openalex updated_date 2026/07/28

Abstract

Character-level models of tokens have been shown to be effective at dealing\nwith within-token noise and out-of-vocabulary words. But these models still\nrely on correct token boundaries. In this paper, we propose a novel end-to-end\ncharacter-level model and demonstrate its effectiveness in multilingual\nsettings and when token boundaries are noisy. Our model is a semi-Markov\nconditional random field with neural networks for character and segment\nrepresentation. It requires no tokenizer. The model matches state-of-the-art\nbaselines for various languages and significantly outperforms them on a noisy\nEnglish version of a part-of-speech tagging benchmark dataset. Our code and the\nnoisy dataset are publicly available at http://cistern.cis.lmu.de/semiCRF.\n

Citations

Cited by

Related