vix.ing · top · new · best · stats · spec

How similar are species names and why does this matter for biodiversity data

2026/07/03 by André Menegotto, Cristina Ronquillo, Joaquín Hortal +1 · 2 voices
Agricultural and Biological Sciences · Earth and Planetary Sciences · Environmental Science · #Coleoptera: Cerambycidae studies #Scarabaeidae Beetle Taxonomy and Biogeography #Species Distribution and Climate Change

paper · pdf · doi:10.3897/bdj.14.e196932

openalex publication_date 2026/07/03 · openalex created_date 2026/07/04 · openalex updated_date 2026/07/14

Abstract

Standardising taxonomic names is an essential step in biodiversity studies to ensure robust data aggregation under the most recent accepted species nomenclature. Fuzzy (inexact) matching is widely used in this process to detect correspondences between scientific names that differ due to alternative spelling or orthographic mistakes. Such an approach assumes that species names are sufficiently distinct such that names differing in just a few characters in fact refer to the same taxon, but this has rarely been evaluated. Across c. 230,000 marine species names, we show that name similarity is common: 19.34% of specific epithets differ by three or fewer edits from another epithet within the same genus. Shared epithets are also widespread within and across phyla, occurring in 73% of all marine species; in 6.05% of these cases, the associated genera differ by three or fewer edits. This level of similarity increases the risk of incorrect matches, limiting the reliability of automated text-string tools in biodiversity big data analyses and highlighting the importance of combining post-matching filters with systematic and authorship information in taxonomic workflows to support name resolution beyond orthographic similarity.

Citations

Discussions

Related