vix.ing · top · new · best · stats

Phylogenetic signal in phonotactics

2020/02/29 by Jayden L. Macklin-Cordes, Claire Bowern, Erich R. Round · 153 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · Social Sciences · #Animal Vocal Communication and Behavior #Artificial intelligence #Australian Indigenous Culture and History #Biology #Computer science #Homogeneity (statistics) #Inference #Language and cultural evolution #Lexicon #Linguistics #Machine learning #Phonology #Phonotactics #Phylogenetic network #Phylogenetic tree #Phylogenetics #cs.CL #q-bio.PE

paper · pdf · doi:10.1075/dia.20004.mac

published in Diachronica 38(2), 210-258 (John Benjamins Publishing Company) · Main text: 32 pages, 17 figures, 1 table. Supplementary Information: 17 pages, 1 figure. Code and data available at http://doi.org/10.5281/zenodo.3936353. This article is in review but not yet accepted for publication in a journal

arxiv created 2020/07/14 · openalex publication_date 2020/12/09 · arxiv updated 2021/05/10 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06

Abstract

Abstract Phylogenetic methods have broad potential in linguistics beyond tree inference. Here, we show how a phylogenetic approach opens the possibility of gaining historical insights from entirely new kinds of linguistic data – in this instance, statistical phonotactics. We extract phonotactic data from 112 Pama-Nyungan vocabularies and apply tests for phylogenetic signal, quantifying the degree to which the data reflect phylogenetic history. We test three datasets: (1) binary variables recording the presence or absence of biphones (two-segment sequences) in a lexicon (2) frequencies of transitions between segments, and (3) frequencies of transitions between natural sound classes. Australian languages have been characterized as having a high degree of phonotactic homogeneity. Nevertheless, we detect phylogenetic signal in all datasets. Phylogenetic signal is greater in finer-grained frequency data than in binary data, and greatest in natural-class-based data. These results demonstrate the viability of employing a new source of readily extractable data in historical and comparative linguistics.

Citations