2005/07/13 by William Yang Wang, James W. Minett · 5 citations
Computer Science · Social Sciences · #Natural Language Processing Techniques #Language and cultural evolution #Authorship Attribution and Profiling
paper · doi:10.1111/j.1467-968x.2005.00147.x
openalex publication_date 2005/07/13 · openalex created_date 2025/10/10 · openalex updated_date 2026/06/23
It has been observed that borrowing within a group of genetically related languages often causes the lexical similarities among them to be skewed. Consequently, it has been proposed that borrowing can sometimes be inferred from such skewing. However, heterogeneity in the rate of lexical replacement, as well as borrowing from other languages, can also give rise to skewed lexical similarities. It is important, therefore, to determine to what degree skewing is a statistically significant indicator of borrowing. Here, we describe a statistical hypothesis test for detecting language contact based on skewing of linguistic characters of arbitrary type. Significant probabilities of correct detection of contact are maintained for various contact scenarios, with low false alarm probability. Our experiments show that the test is fairly robust to substantial heterogeneity in the retention rate, both across characters and across lineages, suggesting that the method can provide an objective criterion against which claims of significant skewing due to contact can be tested, pointing the way for more detailed analysis.