2020/06/01 by Edgar Santos–Fernández, Erin E. Peterson, Santos-Fernandez, Edgar +7
Computer Science · Environmental Science · #Applications (stat.AP) #FOS: Computer and information sciences #Geochemistry and Geologic Mapping #Other Statistics (stat.OT) #Species Distribution and Climate Change
paper · pdf · doi:10.48550/arxiv.2006.00741
openalex publication_date 2020/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
Many research domains use data elicited from "citizen scientists" when a\ndirect measure of a process is expensive or infeasible. However, participants\nmay report incorrect estimates or classifications due to their lack of skill.\nWe demonstrate how Bayesian hierarchical models can be used to learn about\nlatent variables of interest, while accounting for the participants' abilities.\nThe model is described in the context of an ecological application that\ninvolves crowdsourced classifications of georeferenced coral-reef images from\nthe Great Barrier Reef, Australia. The latent variable of interest is the\nproportion of coral cover, which is a common indicator of coral reef health.\nThe participants' abilities are expressed in terms of sensitivity and\nspecificity of a correctly classified set of points on the images. The model\nalso incorporates a spatial component, which allows prediction of the latent\nvariable in locations that have not been surveyed. We show that the model\noutperforms traditional weighted-regression approaches used to account for\nuncertainty in citizen science data. Our approach produces more accurate\nregression coefficients and provides a better characterization of the latent\nprocess of interest. This new method is implemented in the probabilistic\nprogramming language Stan and can be applied to a wide number of problems that\nrely on uncertain citizen science data.\n