2026/01/23 by Zachary M. Portman, Bethanne Bruninga‐Socolar, Marissa H. Chase +4 · 1 voice
Agricultural and Biological Sciences · Biochemistry, Genetics and Molecular Biology · #Insect and Arachnid Ecology and Behavior #Insect and Pesticide Research #Plant and animal studies
paper · doi:10.1093/aesa/saag002
openalex publication_date 2026/01/23 · openalex created_date 2026/02/15 · openalex updated_date 2026/07/30
We thank Cobb et al. (2026) for continuing the discussion on the important issue of ghost records and unverifiable bee specimen data in the United States (Portman et al. 2025). Unfortunately, Cobb et al. largely focus on irrelevant details and semantics—particularly the name of the lab, occurrence data versus monitoring data, the definition of a preserved specimen—while failing to address any of the core points of our paper. Ghost records, or records that are reported as specimen records but lack a physical specimen, are not verifiable and can lead to incorrect conclusions about monitoring and a variety of other uses, including assessing trends, mapping distributions, or measuring biodiversity. We examined the issue of ghost records in Portman et al. (2025) and found that: An estimated 25%–28% of all records in the datasets we examined have had their species identifications invalidated due to taxonomic changes such as the description of cryptic species, and these species identifications can only be corrected by reexamination of the specimens. There is currently no way for scientists using the data to distinguish between a “preserved specimen” that has been preserved and a “preserved specimen” that has been destroyed. Numerous species in the datasets we examined show patterns that are consistent with taxonomic artifacts, such as species that only occur in large numbers in the data after they have been newly described. Cobb et al. claim that we “misname” the USGS Bee Inventory and Monitoring Lab (BIML) and claim the name was changed to the Bee Lab (BL) in 2022. However, while the BL/BIML has used several names over the past few years, it has continued to widely use the Bee Inventory and Monitoring Lab name and the BIML acronym. For example, the lab website called it “The Bee Inventory and Monitoring Lab” up until February 2025 (Eastern Ecological Science Center 2025). In addition, a circular published by the USGS in June 2025, which includes the head of the BL/BIML as an author, referred to the lab as both the “USGS Interagency Native Bee Lab” and “The USGS Native Bee Inventory and Monitoring Lab (BIML)” (Otto et al. 2025a), though most instances of the BIML name and acronym were subsequently removed in a post-publication revision (Otto et al. 2025b). The changes to the website and circular occurred only after our original article had been reviewed and accepted, respectively. Cobb et al. claim that the “USGS Bee Lab Does Not Conduct Monitoring” and “In our view referring to the Bee Lab data as monitoring data is inaccurate.” However, our language is consistent with the language used in the 2 main papers we analyzed by Kammerer et al. (2020) and Collado et al. (2019), which refer to the data as “one of the few existing long-term monitoring datasets for wild bees in the United States” and “a dataset from an extensive monitoring programme for bees in the north-east and midwest United States”, respectively. Further, the BL/BIML data are described and used in the peer-reviewed literature as a monitoring dataset or the product of a monitoring program, including in multiple articles published by authors of the Cobb et al. letter. For example: Kammerer et al. (2020): “Our goal here was to organize, validate, and share an analysis-ready version of one of the few existing long-term, wild-bee monitoring datasets in the United States. Since 1999, the Native Bee Inventory and Monitoring Lab (BIML) of the United States Geological Survey has collected, curated, and identified over 99,000 specimens of more than 314 wild bee species from over 1400 sites in Maryland, Delaware, and Washington DC.” Kammerer et al. (2021): “Before 2003, while developing their monitoring program, BIML employed a variety of sampling methods.” Quinlan, Doser, Kammerer, and Grozinger (2024): “To do so, we leveraged a wild bee monitoring dataset, curated by USGS and USFWS (Droege and Maffei, 2023)” Rousseau, Woodard, Jepsen, Du Clos, et al. (2024): “We also note that some major efforts to collect bee data across the US, such as through the USGS Native Bee Inventory and Monitoring Program, have only recently uploaded a complete version of their records” We find it disappointing that Cobb et al. imply we misrepresented data from the BL/BIML, yet make no attempt to explain their own contradictory published statements on this topic. Finally, Cobb et al. make the semantic distinction that the BL/BIML dataset is composed of “occurrence data” rather than “monitoring data.” Cobb et al. further state: “the data are unambiguously non-standardized, presence-only occurrence records, and should not be interpreted as monitoring records where there is any expectation of tracking changes in the relative abundance of species across years”. However, this is contradicted by the public Global Biodiversity Information Facility (GBIF) page where the BL/BIML data are described, which states: “Number of times a species is recorded in this dataset does not represent actual species abundance or common-ness but does offer an indication of fluctuations in population size” (Droege and Maffei 2025). What are monitoring data intended for if not “an indication of fluctuations in population size”? We believe it is an overly-fine semantic distinction to state that the data “offer an indication of fluctuations in population size” but are not monitoring data. In Portman et al. (2025), we coined the term “ghost records” to refer to specimen records that exist in databases but not in reality. We noted that in the literature, a “preserved specimen” has traditionally been assumed to be a specimen permanently deposited in a museum collection and not one that has been destroyed or dispersed without record. We highlighted how the use of the term “preserved specimen” to refer to destroyed specimens is an issue because it causes confusion and it means users of the data cannot differentiate between a “preserved specimen” that has been preserved versus a “preserved specimen” that has been destroyed. Scientists should be able to assess the quality of the data they are using. We believe that if existing terms or data standards do not adequately convey necessary information (such as whether a specimen has been destroyed or not), that is an issue to be addressed and solved, not argued away. Cobb et al. do not address these issues other than to point out that “There is no expectation of vouchering in perpetuity when a dwc: basisOfRecord is a ‘preserved specimen’; the term is appropriately applied when the original record was based on the examination of a collected specimen.” Cobb et al. do not address our points about how this creates confusion and prevents accurate communication in cases where a “preserved specimen” has been deliberately destroyed after it has been examined. In Portman et al. (2025), we were unable to obtain information on how many specimens from the BL/BIML dataset should be considered ghost records. Cobb et al. provide some numbers that can help fill in those gaps: a synoptic collection at the BL/BIML, an estimated 7,000 specimens deposited at the Smithsonian, and 19,000 specimens identified at other natural history collections. Currently, the BL/BIML data on GBIF report approximately 515,000 specimen records of bees (GBIF.org 2025), which leaves nearly 500,000 bee specimens unaccounted for. We believe it is reasonable to estimate that over 400,000 BL/BIML bee specimens have been destroyed and should be considered ghost records. This must remain an estimate because all of the BIML records continue to be reported as “preserved specimens” and the collection code as “BIML” (GBIF 2025), regardless of whether the specimens have been destroyed or dispersed. Cobb et al. take issue with us focusing on data from BL/BIML. However, data from the BL/BIML are impossible to ignore because it is the single largest source of recent bee specimen records in the United States. The data portal GBIF report approximately 1.8 million bee specimen records in the United States since the year 2000 (GBIF.org 2025), and 515,000 of those records are sourced from the BL/BIML dataset. As a result, the estimated 400,000+ ghost records represent a major portion of publicly available bee specimen data in the United States. The number and proportion of collected specimens that should be retained and deposited in museum collections can be reasonably discussed and debated. However, this can only be done if recommendations are accurately conveyed. In their letter, Cobb et al. state “Portman et al. refer to Packer et al. (2018) specimen retention recommendations; however, Packer et al. (2018) only recommend vouchering focal species exemplars, as is common practice globally.” This is a major misrepresentation of Packer et al. (2018). Indeed, many of the claims made by Cobb et al., including their emphasis on resource limitations and their conflation of invalidated and misidentified specimens, have already been addressed by Packer et al. (2018), which we strongly recommend reading and is worth quoting at length here: [Recommendation] iii) While vouchering of specimens has long been recommended as standard practice (e.g. Huber 1998), whenever possible, all specimens included in larger scale survey research should be housed in a named repository. The obvious reason for this is that misidentifications are likely to occur not only with the vouchered exemplar chosen to represent the identity of that named species, but also in the non-vouchered material: species X may turn out to be a mixture of species X, Y, and Z. When taxon concepts change as a result of more recently uncovered taxonomic complexity, reinterpretations of earlier data will be required and this will not be possible without all material being available for future research. We have heard it said that institutional repositories would not have the facilities to maintain the specimens that would result if this recommendation were implemented. But this seems to us as a circular argument—the more these facilities are used the greater their perceived value which should result in increased allocation of resources. Furthermore, arguments for maintaining, or increasing the resources available to such specimen housing institutions would be strengthened by statements such as Whats in a Name?’. Scientists are increasingly encouraged to make their data available to posterity through archiving (e.g. Whitlock et al. 2010). All this effort will count for little if there is no way of validating the names of the taxa whose data are stored. Cobb et al. do not directly address the issue of whether ghost records can be fixed, but they propose “several options for adjusting existing records to be consistent with accepted taxonomic concepts and/or resolve the true distribution of newly delineated taxa.” They propose that researchers could (i) exclude the affected taxa from analyses, (ii) examine the records at the level of species-group or genus rather than species, (iii) examine species from other insect collections, and (iv) re-sampling for taxa of interest. All of the solutions by Cobb et al. boil down to excluding ghost records from analyses. We disagree that this represents a valid fix because the species-level identifications of destroyed specimens cannot be recovered or verified. Further, resampling in many of these areas may not be feasible, particularly for the over 70,000 BL/BIML records that have “National Park” in the locality (GBIF.org 2025). Kammerer et al. (2020) state: “Our goal here was to organize, validate, and share an analysis-ready version of one of the few existing long-term, wild-bee monitoring datasets in the United States.” We analyzed that dataset. We found some unfixable taxonomic issues and demonstrated that those issues are also present in the broader GBIF dataset, with important ramifications for anyone analyzing publicly available bee specimen data. The confusion and disagreement over such basic questions as “what is a preserved specimen?” and “are the BL/BIML data monitoring data?” show how unclear appropriate use is and leaves our field open to erroneous analyses and conclusions about bee diversity and trends. We agree with Cobb et al. that large amounts of data are needed to address some important scientific questions. However, the quantity of specimen data only matters if there is some minimum quality standard. If data cannot be verified, a cornerstone of the scientific process, then it is not useful for answering scientific questions, regardless of whether it is monitoring data or occurrence data. With unverifiable data making up a large portion of bee specimen data in the United States, this issue feeds into the broader replication crisis and becomes an issue that needs to be addressed and reckoned with by the entire field. Disagreements, critiques, and debates are an integral part of the scientific process. We thank Cobb et al. for contributing to the discussion of this important issue. None declared. None declared.