vix.ing · top · new · best · stats · spec

Noise, bias and data limitations: Virtual benchmarking of species distribution models' robustness for prediction and inference

2026/07/20 by Emma Chollet Ramampiandra, Gaspard Fragnière, Andreas Scheidegger +1 · 1 voice
Environmental Science · #Environmental DNA in Biodiversity Studies #Freshwater macroinvertebrate diversity and ecology #Species Distribution and Climate Change

paper · doi:10.1016/j.ecoinf.2026.103941

openalex publication_date 2026/07/20 · openalex created_date 2026/07/21 · openalex updated_date 2026/07/30

Abstract

Species distribution models (SDMs) are widely used to describe and predict species occurrences from environmental data, but calibration datasets are typically affected by noise and bias. We test the robustness of SDMs to common data limitations by generating ecologically realistic virtual benchmark data with a mechanistic food-web model and then systematically degrading the data quality. These datasets are used to train SDMs of different complexities, which we evaluate not only on their predictive performance, but also on their degree of overfitting and their ability to recover the true environmental response shapes. While SDMs explicitly only capture environmental responses, species occurrences in nature are also affected by dispersal limitation and biotic interactions. To account for this, we generated virtual benchmark data with the model Streambugs, which explicitly includes biotic interactions and dispersal limitation at the catchment scale, to simulate presence absence data and responses of 129 stream macroinvertebrate taxa to eight environmental factors. We then systematically applied four common types of data-limitation scenarios: reduced sample size, missing predictors, random noise added to one predictor, and bias in the species detection process. Each type was assessed at four levels of degradation, and we additionally evaluated combined scenarios including all limitations. We compared statistical and machine-learning SDMs covering a range of model complexities for 35 taxa. More flexible models achieved the highest predictive performance under the best-case scenario but were sensitive to data degradation, exhibiting strong overfitting that leads to irregular response shapes, which did not reflect the data generating processes. In contrast, simpler models produced more consistent predictions and response shapes across scenarios. We illustrate how SDMs differ in their robustness to common data limitations and reveal trade-offs between predictive performance, overfitting and ecological interpretability, enabling informed decisions about model complexity according to data quality and study objectives.

Citations

Discussions

Related