Calibration ofρValues for Testing Precise Null Hypotheses
2001/02/01 by Thomas Sellke, M. J Bayarri, M. J. Bayarri +2 · 862 citations
Decision Sciences · Mathematics · #Alternative hypothesis #Bayesian inference #Bayesian probability #Calibration #Conditional probability #Econometrics #Frequentist inference #Frequentist probability #Mathematics #Meta-analysis and systematic reviews #Null hypothesis #Odds #Posterior probability #Statistical Methods in Clinical Trials #Statistical hypothesis testing #Statistical power #Statistics #p-value
paper · doi:10.1198/000313001300339950
published in The American Statistician 55(1), 62-71 (Taylor & Francis)
openalex publication_date 2001/02/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/04
Abstract
P values are the most commonly used tool to measure evidence against a hypothesis or hypothesized model. Unfortunately, they are often incorrectly viewed as an error probability for rejection of the hypothesis or, even worse, as the posterior probability that the hypothesis is true. The fact that these interpretations can be completely misleading when testing precise hypotheses is first reviewed, through consideration of two revealing simulations. Then two calibrations of a ρ value are developed, the first being interpretable as odds and the second as either a (conditional) frequentist error probability or as the posterior probability of the hypothesis.
Cited by
- How the Maximal Evidence ofP-Values Against Point Null Hypotheses Depends on Sample Size
- Null Hypothesis Significance Testing Interpreted and Calibrated by Estimating Probabilities of Sign Errors: A Bayes-Frequentist Continuum
- Insufficient evidence for DMS and DMDS in the atmosphere of K2-18 b
- Valid P -Values Behave Exactly as They Should: Some Misleading Criticisms of P -Values and Their Resolution With S -Values
- The virtues of frugality — why cosmological observers should release their data slowly
- A Framework for Validation of Computer Models
- Admissible ways of merging p-values under arbitrary dependence
- Atmospheric retrieval evidence for water isotopologue HDO on exoplanet WASP-39b
- A Systematic Search for Trace Molecules in the Atmosphere of Exoplanet K2-18 b
- Quantitative mobile gamma-ray spectrometry through Bayesian inference
- Bayesian Methods in Cosmology
- Separating flare and secondary atmospheric signals with RADYN modeling of near-infrared JWST transmission spectroscopy observations of TRAPPIST-1
- The atmospheric composition of TOI-270 d
- A fast Monte Carlo test for preferential sampling
- A Bayesian Perspective on Evidence for Evolving Dark Energy
- Rejection Odds and Rejection Ratios: A Proposal for Statistical Practice in Testing Hypotheses
- Is There a Free Lunch in Inference?
- Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations
- A small-sample Bayesian information criterion that does not overstate the evidence, with an application to calibrating p-values from likelihood-ratio tests
- Default Bayes factors for ANOVA designs
- AIC model selection and multimodel inference in behavioral ecology: some background, observations, and comparisons
- A practical solution to the pervasive problems ofp values
- Bayesian inference for psychology. Part I: Theoretical advantages and practical ramifications
- Divergence versus decision P ‐values: A distinction worth making in theory and keeping in practice: Or, how divergence P ‐values measure evidence even when decision P ‐values do not
- Invited Commentary: The Need for Cognitive Science in Methodology
- Bayesian Model Comparison and Significance: Widespread Errors and how to Correct Them
- Calibrating p‐values in ecology: a practical framework for integrating prior plausibility into statistical inference
- Bayes factors vs. P-values
- BFpack: Flexible Bayes Factor Testing of Scientific Theories in R
- On the predicted replicability of two decades of experimental research on system justification: A Z‐curve analysis
- Looking guilty: Handcuffing suspects influences judgements of deception
- JWST-TST DREAMS: Sulfur dioxide in the atmosphere of the Neptune-mass planet HAT-P-26 b from NIRSpec G395H transmission spectroscopy
- Reverse‐Bayes methods for evidence assessment and research synthesis
- Hydrocarbon Hazes on Temperate sub-Neptune K2-18b supported by data from the James Webb Space Telescope
- Combining p-values via averaging
- To Vary or Not To Vary: A Flexible Empirical Bayes Factor for Testing Variance Components
- Is the p-value a good measure of evidence? An asymptotic consistency criterion
- There are natural scores: Full comment on Shafer, "Testing by betting: A strategy for statistical and scientific communication"
- Profiles of impulsivity and alcohol use: Unveiling personality, cognitive traits, and DSM diagnoses
- The first large absorption survey in H i (FLASH): II. Pilot survey data release and first results
- Technical Issues in the Interpretation of S-values and Their Relation to Other Information Measures
- Rejection odds and rejection ratios: A proposal for statistical practice in testing hypotheses
- Null hypothesis significance tests: A mix-up of two different theories, the basis for widespread confusion and numerous misinterpretations
- Resurrecting the One-Sided P-value as a Likelihood Ratio
- From conformal to probabilistic prediction
- An Assessment of Organics Detection and Characterization on the Surface of Europa with Infrared Spectroscopy
- SiO and a super-stellar C/O ratio in the atmosphere of the giant exoplanet WASP-121b
- Blending Bayesian and frequentist methods according to the precision of prior information with an application to hypothesis testing
- On the Ubiquity of Information Inconsistency for Conjugate Priors
- A note on data splitting with e-values: online appendix to my comment on Glenn Shafer's "Testing by betting"
- Testing Differences Between Pathogen Compositions with Small Samples and Sparse Data
- Current Status of Open Science and Statistical Analysis in <i>The Japanese Journal of Educational Psychology</i>:
- Constraint-based Causal Discovery from Multiple Interventions over Overlapping Variable Sets
- Bayesian Model for Multiple Change-points Detection in Multivariate Time Series
- Evaluation of three methods for calculating statistical significance when incorporating a systematic uncertainty into a test of the background-only hypothesis for a Poisson process
- p-values for model evaluation
- A BAYESIAN APPROACH TO COMPARING COSMIC RAY ENERGY SPECTRA
- The preregistration revolution
- Why P Values Are Not a Useful Measure of Evidence in Statistical Significance Testing
- Confusion Over Measures of Evidence (p's) Versus Errors (α's) in Classical Statistical Testing
- On posterior probability and significance level: application to the power spectrum of HD 49 933 observed by CoRoT
- Frequentist tests for Bayesian models
- Test Martingales, Bayes Factors and p-Values
- Panchromatic view of the frigid Jovian exoplanet COCONUTS-2 b
- Has DESI detected exponential quintessence?
- Global analysis of aberrant pre-mRNA splicing in glioblastoma using exon expression arrays. [europepmc]
- A nomogram for P values. [europepmc]
- Dietary intake of micronutrients and the risk of developing bladder cancer: results from the Belgian case-control study on bladder cancer risk. [europepmc]
- Significance testing as perverse probabilistic reasoning. [europepmc]
- A Bayesian hierarchical diffusion model decomposition of performance in Approach-Avoidance Tasks. [europepmc]
- An investigation of the false discovery rate and the misinterpretation of p-values. [europepmc]
- Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. [europepmc]
- A Tutorial on Hunting Statistical Significance by Chasing N . [europepmc]
- Integrative analysis of genetic data sets reveals a shared innate immune component in autism spectrum disorder and its co-morbidities. [europepmc]
- Bayesian techniques for analyzing group differences in the Iowa Gambling Task: A case study of intuitive and deliberate decision-makers. [europepmc]
- The earth is flat ( p > 0.05): significance thresholds and the crisis of unreplicable research. [europepmc]
- Bayesian inference for psychology. Part I: Theoretical advantages and practical ramifications. [europepmc]
- When Null Hypothesis Significance Testing Is Unsuitable for Research: A Reassessment. [europepmc]
- Daily rhythms and enrichment patterns in the transcriptome of the behavior-manipulating parasite Ophiocordyceps kimflemingiae. [europepmc]
- The reproducibility of research and the misinterpretation of p -values. [europepmc]
- Lower rotational inertia and larger leg muscles indicate more rapid turns in tyrannosaurids than in other large theropods. [europepmc]
- Cerebellar Repetitive Transcranial Magnetic Stimulation and Noisy Galvanic Vestibular Stimulation Change Vestibulospinal Function. [europepmc]
- Selective Enhancement of Object Representations through Multisensory Integration. [europepmc]
- Semantic and cognitive tools to aid statistical science: replace confidence and significance by compatibility and surprise. [europepmc]
- Hyperspectral imaging and robust statistics in non-melanoma skin cancer analysis. [europepmc]
- Assessing treatment effects and publication bias across different specialties in medicine: a meta-epidemiological study. [europepmc]
- Heavy Metal Levels in Milk and Cheese Produced in the Kvemo Kartli Region, Georgia. [europepmc]
- Persistence of information flow: A multiscale characterization of human brain. [europepmc]
- The Effects of a Non-Technical Skills Training Program on Emotional Intelligence and Resilience in Undergraduate Nursing Students. [europepmc]
- Biosynthesis of bacterial cellulose nanofibrils in black tea media by a symbiotic culture of bacteria and yeast isolated from commercial kombucha beverage. [europepmc]
Related