2026/04/27 by André Menegotto, Cristina Ronquillo, Joaquín Hortal +1 · 1 voice
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · Environmental Science · #Biomedical Text Mining and Ontologies #Coleoptera: Cerambycidae studies #Species Distribution and Climate Change
paper · pdf · doi:10.3897/arphapreprints.e196971
openalex publication_date 2026/04/27 · openalex created_date 2026/04/28 · openalex updated_date 2026/08/02
Standardising taxonomic names is an essential step in biodiversity studies to ensure robust data aggregation and up-to-date, valid species nomenclature. Fuzzy (inexact) matching is widely used in this process to detect correspondences between scientific names that differ due to typographical errors. Such an approach assumes that species names are sufficiently distinct such that names differing in just a few characters in fact refer to the same taxon, but this has rarely been evaluated. Across c. 230,000 marine species names, we show that name similarity is common: 28.37% of specific epithets differ by three or fewer edits from another epithet within the same genus. Shared epithets are also widespread within and across phyla, occurring in 73% of all marine species; in 7.35% of these cases, the associated genera differ by three or fewer edits. This level of similarity increases the risk of incorrect matches, limiting the reliability of automated text-string tools in biodiversity big data analyses and highlighting the importance of considering systematic and authorship information into taxonomic workflows to support name resolution beyond orthographic similarity.