1998/09/28 by Luis Alejandro Márquez–Martínez, L. Marquez, L. Padro +6
Computer Science · #Algorithms and Data Compression #Computation and Language (cs.CL) #FOS: Computer and information sciences #I.2.7 #Natural Language Processing Techniques #Speech and dialogue systems #cs.CL
paper · pdf · doi:10.48550/arxiv.cs/9809113
Appears in proceedings of NLP+IA/TAL+AI'98. Moncton, New Brunswick, Canada, 1998
arxiv created 1998/09/28 · openalex publication_date 1998/09/28 · arxiv updated 2009/11/30 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
We present a bootstrapping method to develop an annotated corpus, which is specially useful for languages with few available resources. The method is being applied to develop a corpus of Spanish of over 5Mw. The method consists on taking advantage of the collaboration of two different POS taggers. The cases in which both taggers agree present a higher accuracy and are used to retrain the taggers.