2026/03/17 by Katharina F. Heil, Tyler Alioto, Astrid Böhne +18 · 1 voice
Computer Science · Environmental Science · #Research Data Management Practices #Environmental DNA in Biodiversity Studies #Species Distribution and Climate Change
paper · pdf · doi:10.3897/rio.12.e187033
openalex publication_date 2026/03/17 · openalex created_date 2026/03/18 · openalex updated_date 2026/07/02
Biodiversity genomics is converging from historically separate approaches — DNA barcoding and reference genome sequencing — into an integrated digital ecosystem driven by shared data stewardship principles: transparent provenance, persistent identifiers and interoperable repositories. We demonstrate how these workflows can operate within a unified informatics architecture spanning data generation, validation, publication and reuse. We describe coordinated infrastructure, including the European BOLD mirror, ERGA Genome Tracking Console and metadata platforms COPO and PlutoF. These systems employ harmonised validation pipelines, shared metadata standards that bridge the Darwin Core and Genomic Standards Consortium vocabularies and automated data exchange amongst ENA, UNITE and GBIF. Workflows in Galaxy, Nextflow and Snakemake are registered in WorkflowHub as Research Object Crates (RO-Crates), ensuring reproducibility and complete provenance. Key outcomes include comprehensive data flow documentation, automated quality control using BUSCO and ERGA Assembly Reports and robust specimen-to-data linkage. We identify challenges in metadata harmonisation, distributed tracking, collaborative attribution and infrastructure sustainability and provide recommendations for addressing them through existing platforms and emerging RO-Crate standards. This work establishes practical foundations for treating biodiversity molecular data as a continuum, demonstrating how FAIR principles can scale to continental initiatives.