2011/11/01 by Dennis D. Boos, Leonard A. Stefanski · 170 citations
Computer Science · Decision Sciences · Mathematics · #Artificial intelligence #Computer science #Context (archaeology) #Contrast (vision) #Data Analysis with R #Econometrics #Geography #Mathematics #Meta-analysis and systematic reviews #Multiple comparisons problem #Null hypothesis #Replicate #Reproducibility #Rounding #Sample (material) #Sample size determination #Statistical Methods in Clinical Trials #Statistical hypothesis testing #Statistical power #Statistical significance #Statistics #Value (mathematics) #p-value
paper · pdf · doi:10.1198/tas.2011.10129
published in The American Statistician 65(4), 213-221 (Taylor & Francis)
openalex publication_date 2011/11/01 · openalex created_date 2016/06/24 · openalex updated_date 2026/08/05
P-values are useful statistical measures of evidence against a null hypothesis. In contrast to other statistical estimates, however, their sample-to-sample variability is usually not considered or estimated, and therefore not fully appreciated. Via a systematic study of log-scale p-value standard errors, bootstrap prediction bounds, and reproducibility probabilities for future replicate p-values, we show that p-values exhibit surprisingly large variability in typical data situations. In addition to providing context to discussions about the failure of statistical results to replicate, our findings shed light on the relative value of exact p-values vis-a-vis approximate p-values, and indicate that the use of *, **, and *** to denote levels .05, .01, and .001 of statistical significance in subject-matter journals is about the right level of precision for reporting p-values when judged by widely accepted rules for rounding statistical estimates.