Valid P -Values Behave Exactly as They Should: Some Misleading Criticisms of P -Values and Their Resolution With S -Values
2019/03/20 by Sander Greenland · 43 citations
Computer Science · Decision Sciences · #Rough Sets and Fuzzy Logic #Multi-Criteria Decision Making #Bayesian Modeling and Causal Inference
paper · pdf · doi:10.1080/00031305.2018.1529625
Abstract
The present note explores sources of misplaced criticisms of P-values, such as conflicting definitions of “significance levels” and “P-values” in authoritative sources, and the consequent misinterpretation of P-values as error probabilities. It then discusses several properties of P-values that have been presented as fatal flaws: That P-values exhibit extreme variation across samples (and thus are “unreliable”), confound effect size with sample size, are sensitive to sample size, and depend on investigator sampling intentions. These properties are often criticized from a likelihood or Bayesian framework, yet they are exactly the properties P-values should exhibit when they are constructed and interpreted correctly within their originating framework. Other common criticisms are that P-values force users to focus on irrelevant hypotheses and overstate evidence against those hypotheses. These problems are not however properties of P-values but are faults of researchers who focus on null hypotheses and overstate evidence based on misperceptions that p = 0.05 represents enough evidence to reject hypotheses. Those problems are easily seen without use of Bayesian concepts by translating the observed P-value p into the Shannon information (S-value or surprisal) –log2(p).
Citations
Cited by
- Moving to a World Beyond “ p < 0.05”
- Inferential Statistics as Descriptive Statistics: There Is No Replication Crisis if We Don’t Expect Replication
- Giving less power to statistical power
- From means to meaning in the study of sex/gender differences and similarities
- From significant to meaningful: ATOMizing the study of sex differences and similarities
- Why and how we should join the shift from significance testing to estimation
- Causal clarity in statistical software
- A small-sample Bayesian information criterion that does not overstate the evidence, with an application to calibrating p-values from likelihood-ratio tests
- Artificial intelligence and dichotomania
- Divergence versus decision<i>P</i>‐values: A distinction worth making in theory and keeping in practice: Or, how divergence<i>P</i>‐values measure evidence even when decision<i>P</i>‐values do not
- Interpreting statistical evidence with S values: an illustration with trials of goal‐directed haemodynamic therapy
- <i>P</i> value functions: An underused method to present research results and to promote quantitative reasoning
- The role of multiple global change factors in driving soil functions and microbial biodiversity
- A tale of two scripts: Applying the principle of least complexity to simplified and traditional Chinese
- Assessing replication success: a taxonomy and systematic comparison of replication success measures for prospective replications
- The Search for Truth through Data: NP Decision Processes, ROC Functions, P-Functionals, Knowledge Updating and Sequential Learning
- There are natural scores: Full comment on Shafer, "Testing by betting: A strategy for statistical and scientific communication"
- Common misinterpretations of statistical significance and P-values in dairy research
- Multiple comparisons controversies are about context and costs, not frequentism versus Bayesianism. [europepmc]
- Systematic review of the use of "magnitude-based inference" in sports science and medicine. [europepmc]
- Improving practices and inferences in developmental cognitive neuroscience. [europepmc]
- Semantic and cognitive tools to aid statistical science: replace confidence and significance by compatibility and surprise. [europepmc]
- Comparison of Plant Metabolites in Root Exudates of Lolium perenne Infected with Different Strains of the Fungal Endophyte Epichloë festucae var. lolii . [europepmc]
- There is life beyond the statistical significance. [europepmc]
- Examining the Effect of Context, Beliefs, and Values on UK Farm Veterinarians' Antimicrobial Prescribing: A Randomized Experimental Vignette and Cross-Sectional Survey. [europepmc]
- Preserved structural connectivity mediates the clinical effect of thrombolysis in patients with anterior-circulation stroke. [europepmc]
- Scan Once, Analyse Many: Using Large Open-Access Neuroimaging Datasets to Understand the Brain. [europepmc]
- The role of plant labile carbohydrates and nitrogen on wheat-aphid relations. [europepmc]
- Australian Lentil Breeding Between 1988 and 2019 Has Delivered Greater Yield Gain Under Stress Than Under High-Yield Conditions. [europepmc]
- Use of the p-values as a size-dependent function to address practical differences when analyzing large datasets. [europepmc]
- Anterograde interference emerges along a gradient as a function of task similarity: A behavioural study. [europepmc]
- Providing Evidence for the Null Hypothesis in Functional Magnetic Resonance Imaging Using Group-Level Bayesian Inference. [europepmc]
- Assessing Knowledge, Beliefs, and Behaviors around Antibiotic Usage and Antibiotic Resistance among UK Veterinary Students: A Multi-Site, Cross-Sectional Survey. [europepmc]
- Statistical significance and its critics: practicing damaging science, or damaging scientific practice? [europepmc]
- Why and how we should join the shift from significance testing to estimation. [europepmc]
- Replacing statistical significance and non-significance with better approaches to sampling uncertainty. [europepmc]
- Impact of mobile health on maternal and child health service utilization and continuum of care in Northern Ghana. [europepmc]
- Seeing the Error in My " Bayes ": A Quantified Degree of Belief Change Correlates with Children's Pupillary Surprise Responses Following Explicit Predictions. [europepmc]
- Submaximal Fitness Test in Team Sports: A Systematic Review and Meta-Analysis of Exercise Heart Rate Measurement Properties. [europepmc]
- The Acute Demands of Repeated-Sprint Training on Physiological, Neuromuscular, Perceptual and Performance Outcomes in Team Sport Athletes: A Systematic Review and Meta-analysis. [europepmc]
- Interpreting Randomized Controlled Trials. [europepmc]
- For a proper use of frequentist inferential statistics in public health. [europepmc]
- Common mistakes in biostatistics. [europepmc]
Related