vix.ing · top · new · best · stats · spec

Data interference

2026/06/23 by Matteo Di Cristofaro · 1 voice
Arts and Humanities · Computer Science · Psychology · #Digital Communication and Language #Discourse Analysis in Language Studies #Language, Metaphor, and Cognition

paper · doi:10.1075/ijcl.24116.dic

openalex publication_date 2026/06/23 · openalex created_date 2026/06/24 · openalex updated_date 2026/07/31

Abstract

Abstract Tokenisation is a crucial step for corpus linguistics, as it provides the basis for any applicable quantitative method (e.g. collocations) while ensuring the reliability of qualitative approaches. This paper examines how discrepancies in tokenisation affect the representation of language data and the validity of analytical findings. Investigating the challenges posed by emojis and homoglyphs, the study highlights the necessity of pre-processing these elements to maintain corpus fidelity to the source data. The research presents methods for ensuring that digital texts are accurately represented in corpora, thereby supporting reliable linguistic analysis and facilitating the repeatability of linguistic interpretations. The findings emphasise the necessity of a detailed understanding of both linguistic and technical aspects involved in digital textual data to enhance the accuracy of corpus analysis, and have significant implications for both quantitative and qualitative approaches in corpus-based research.

Citations

Discussions

Related