vix.ing · top · new · best · stats · spec

A Portuguese Native Language Identification Dataset

2018/04/30 by Iria del Río, del Río, Iria, Marcos Zampieri +3
Computer Science · Health Professions · #Authorship Attribution and Profiling #Computation and Language (cs.CL) #FOS: Computer and information sciences #Interpreting and Communication in Healthcare #Natural Language Processing Techniques #cs.CL

paper · pdf · doi:10.48550/arxiv.1804.11346

Proceedings of The 13th Workshop on Innovative Use of NLP for Building Educational Applications (BEA)

arxiv created 2018/04/30 · openalex publication_date 2018/04/30 · arxiv updated 2018/05/01 · openalex created_date 2022/08/15 · openalex updated_date 2026/07/28

Abstract

In this paper we present NLI-PT, the first Portuguese dataset compiled for Native Language Identification (NLI), the task of identifying an author's first language based on their second language writing. The dataset includes 1,868 student essays written by learners of European Portuguese, native speakers of the following L1s: Chinese, English, Spanish, German, Russian, French, Japanese, Italian, Dutch, Tetum, Arabic, Polish, Korean, Romanian, and Swedish. NLI-PT includes the original student text and four different types of annotation: POS, fine-grained POS, constituency parses, and dependency parses. NLI-PT can be used not only in NLI but also in research on several topics in the field of Second Language Acquisition and educational NLP. We discuss possible applications of this dataset and present the results obtained for the first lexical baseline system for Portuguese NLI.

Related