2025/12/29 by Marcelo Finger, Maria Clara Paixão de Sousa, Cristiane NAMIUTI +7 · 1 voice
Arts and Humanities · Computer Science · Social Sciences · #Digital Humanities and Scholarship #Language and cultural evolution #Natural Language Processing Techniques
paper · pdf · doi:10.25189/2675-4916.2025.v6.n4.id812
openalex publication_date 2025/12/29 · openalex created_date 2025/12/30 · openalex updated_date 2026/08/03
This paper presents the challenges of building Carolina, a large open corpus of Brazilian Portuguese texts developed since 2020 using the Web as Corpus methodology enhanced with concerns about provenance and typology (WaC-wiPT). The corpus aims to serve both as a reliable source for research in Linguistics and as an important resource for Computer Science research on language models. Above all, this endeavor aims at removing Portuguese from the set of “low-resource languages”. This paper details the construction methodology of Carolina, with special attention to the issue of describing provenance and typology according to international standards, while briefly describing its relationship with other existing corpora, its current state of development, and its future directions.