vix.ing · top · new · best · stats

Why Attention? Analyzing and Remedying BiLSTM Deficiency in Modeling Cross-Context for NER

2019/10/07 by Peng-Hsuan Li, Li, Peng-Hsuan, Tsu-Jui Fu +3 · 1 citation
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling #cs.CL

paper · pdf · doi:10.48550/arxiv.1910.02586

This short article is obsolete, as its content is contained in the full paper arXiv:1908.11046, which will also be published by AAAI 2020

arxiv created 2020/02/29 · arxiv updated 2020/03/03

Abstract

State-of-the-art approaches of NER have used sequence-labeling BiLSTM as a core module. This paper formally shows the limitation of BiLSTM in modeling cross-context patterns. Two types of simple cross-structures -- self-attention and Cross-BiLSTM -- are shown to effectively remedy the problem. On both OntoNotes 5.0 and WNUT 2017, clear and consistent improvements are achieved over bare-bone models, up to 8.7% on some of the multi-token mentions. In-depth analyses across several aspects of the improvements, especially the identification of multi-token mentions, are further given.

Citations

Cited by

Related