2026/07/14 by Petr Plecháč, Artjoms Šeļa · 2 voices
Arts and Humanities · Computer Science · #Artificial Intelligence in Games #Authorship Attribution and Profiling #Boosting (machine learning) #Cluster analysis #Curse of dimensionality #Czech #Digital Humanities and Scholarship #Dimension (graph theory) #Dimensionality reduction #Pairwise comparison #Pattern recognition (psychology) #Sequence (biology)
paper · pdf · doi:10.1371/journal.pone.0340514
published in PLoS ONE 21(7), e0340514 (Public Library of Science)
openalex publication_date 2026/07/14 · openalex created_date 2026/07/15 · openalex updated_date 2026/08/01
Fixed poetic forms such as the sonnet, ottava rima, or terza rima are an important feature of European literary traditions, yet large-scale empirical research on their cross-lingual distribution and evolution has been limited so far. This paper introduces a fully language-independent, unsupervised method for identifying recurrent rhyme-based forms using local sequence alignment. Drawing on 187,719 poems from six European traditions (Czech, English, French, German, Italian, Russian) in the PoeTree collection, we encode rhyme schemes in a compact eight-symbol alphabet and apply the Smith-Waterman algorithm via the Metronome package to compute pairwise distances. Dimensionality reduction (UMAP) and density-based clustering (HDBSCAN) yield 61 clusters, many of which align with known fixed forms. Evaluation against existing Czech and Russian annotations shows strong recall, while supervised classification experiments-both within and across languages-demonstrate that form categories are robustly learnable in the induced vector space. We illustrate the potential of such data for literary research in three showcases: cross-tradition influence in 19th-century Czech poetry, topical affinities of selected forms using multilingual topic modeling, and geographic associations revealed through geonym analysis.