vix.ing · top · new · best · stats · spec

Compiling and Processing Historical and Contemporary Portuguese Corpora

2017/10/02 by Marcos Zampieri, Zampieri, Marcos
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Natural Language Processing Techniques #Text Readability and Simplification #Topic Modeling

paper · pdf · doi:10.48550/arxiv.1710.00803

openalex publication_date 2017/10/02 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This technical report describes the framework used for processing three large Portuguese corpora. Two corpora contain texts from newspapers, one published in Brazil and the other published in Portugal. The third corpus is Colonia, a historical Portuguese collection containing texts written between the 16th and the early 20th century. The report presents pre-processing methods, segmentation, and annotation of the corpora as well as indexing and querying methods. Finally, it presents published research papers using the corpora.

Related