2018/03/22 by Phạm Quang Minh, Minh, Pham Quang Nhat
Computer Science · Decision Sciences · #Computation and Language (cs.CL) #Data Quality and Management #FOS: Computer and information sciences #Natural Language Processing Techniques #Topic Modeling
paper · pdf · doi:10.48550/arxiv.1803.08463
openalex publication_date 2018/03/22 · openalex created_date 2022/10/03 · openalex updated_date 2026/07/28
In this report, we describe our participant named-entity recognition system\nat VLSP 2018 evaluation campaign. We formalized the task as a sequence labeling\nproblem using BIO encoding scheme. We applied a feature-based model which\ncombines word, word-shape features, Brown-cluster-based features, and\nword-embedding-based features. We compare several methods to deal with nested\nentities in the dataset. We showed that combining tags of entities at all\nlevels for training a sequence labeling model (joint-tag model) improved the\naccuracy of nested named-entity recognition.\n