The Difference Between “Significant” and “Not Significant” is not Itself Statistically Significant
2006/10/11 by Andrew Gelman, Hal S. Stern, Hal Stern · 2 voices · 1,238 citations
Mathematics · Physics and Astronomy · Psychology · #Cognitive psychology #Computer science #Data mining #Econometrics #Interpretation (philosophy) #Mathematics #Null (SQL) #Null hypothesis #Paranormal Experiences and Beliefs #Psychology #Quantum Mechanics and Applications #Statistical Mechanics and Entropy #Statistical analysis #Statistical hypothesis testing #Statistical significance #Statistics
paper · doi:10.1198/000313006x152649
published in The American Statistician 60(4), 328-331 (Taylor & Francis)
openalex publication_date 2006/10/11 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Abstract
It is common to summarize statistical comparisons by declarations of statistical significance or nonsignificance. Here we discuss one problem with such declarations, namely that changes in statistical significance are often not themselves statistically significant. By this, we are not merely making the commonplace observation that any particular threshold is arbitrary—for example, only a small change is required to move an estimate from a 5.1% significance level to 4.9%, thus moving it into statistical significance. Rather, we are pointing out that even large changes in significance levels can correspond to small, nonsignificant changes in the underlying quantities.The error we describe is conceptually different from other oft-cited problems—that statistical significance is not the same as practical importance, that dichotomization into significant and nonsignificant results encourages the dismissal of observed differences in favor of the usually less interesting null hypothesis of no difference, and that any particular threshold for declaring significance is arbitrary. We are troubled by all of these concerns and do not intend to minimize their importance. Rather, our goal is to bring attention to this additional error of interpretation. We illustrate with a theoretical example and two applied examples. The ubiquity of this statistical error leads us to suggest that students and practitioners be made more aware that the difference between “significant” and “not significant” is not itself statistically significant.
Cited by
- The ASA Statement on p -Values: Context, Process, and Purpose
- Valid P -Values Behave Exactly as They Should: Some Misleading Criticisms of P -Values and Their Resolution With S -Values
- Moving to a World Beyond “ p < 0.05”
- Inferential Statistics as Descriptive Statistics: There Is No Replication Crisis if We Don’t Expect Replication
- P-Value Precision and Reproducibility
- Absence of Systematic Effects of Internalizing Psychopathology on Learning Under Uncertainty
- Functional and representational differences between bilateral inferior temporal numeral areas
- Carbon majors and the scientific case for climate liability
- The Impact of US Drone Strikes on Terrorism in Pakistan
- An Index and Test of Linear Moderated Mediation
- A Framework for Moving Beyond Computational Reproducibility: Lessons from Three Reproductions of Geographical Analyses of COVID‐19
- Extraversion, Gender, and the Perceived Pleasantness of Politics
- The Protective Effect of Marriage for Survival: A Review and Update
- Cents and shenshibility: The role of reward in talker-specific phonetic recalibration
- Hominin locomotion and evolution in the Late Miocene to Late Pliocene
- No Need to Watch: How the Effects of Partisan Media Can Spread via Interpersonal Discussions
- Experiential Versus Instructional Approaches for Eliciting Metacognitive Awareness in AI-Assisted Learning: A Short-Term Longitudinal Study
- Evaluation-Recording Contamination in Learned Nano-Quadrotor Dynamics: A Fresh-Seed Audit
- The Difference Between "Replicable" and "Not replicable" is not Itself Scientifically Replicable
- Semiotic Complexity and Its Epistemological Implications for Modeling Culture
- From significant to meaningful: ATOMizing the study of sex differences and similarities
- Detecting long-range attraction between migrating cells based on p-value distributions
- Toward New Directions in Human Biology: A Roadmap for Anthropological Causal Inference With Observational Data
- The Paradox of Algorithms and Blame on Public Decision-makers
- Evaluating Authoritarian Performance: Historical Legacies and Contemporary Attitudes in Saudi Arabia
- Populist Publics’ Affinity for Nuclear Weapons? Explaining Opposition to the Nuclear Sharing Weapons in European Populist Publics
- Neurodatascience: Past, Present, and Future
- Primer for Experimental Methods in Organization Theory
- P-curve: A key to the file-drawer.
- Low interspecific variation and no phylogenetic signal in additive genetic variance in wild bird and mammal populations
- Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations
- MYSTERY SHOPPING: KLJUČNI ČIMBENIK USPJEŠNOG BENCHMARKINGA U MALOPRODAJI
- Artificial intelligence and dichotomania
- Test-retest reliability of the Welfare Quality® animal welfare assessment protocol for growing pigs
- Power Over Presence: Women’s Representation in Comprehensive Peace Negotiations and Gender Provision Outcomes
- The Separation Plot: A New Visual Method for Evaluating the Fit of Binary Models
- Critical Quantitative Literacy: An Educational Foundation for Critical Quantitative Research
- Mobile money and the social contract: Experimental evidence from Ghana
- How adverse childhood experiences get under the skin: A systematic review, integration and methodological discussion on threat and reward learning mechanisms
- P value functions: An underused method to present research results and to promote quantitative reasoning
- How can we make sound replication decisions?
- Examining the replicability of online experiments selected by a decision market
- Tempered Expectations: A Tutorial for Calculating and Interpreting Prediction Intervals in the Context of Replications
- The Longitudinal Association between Self–Esteem and Depressive Symptoms in Adolescents: Separating Between–Person Effects from Within–Person Effects
- Cardiac timing effects on response speed are modulated by blood pressure but not heart rate variability in healthy young adults
- Still Instrumentally Inclusive
- Metrological uncertainty and the value of informative nulls in marketing research
- A multiverse lack of replication for working memory capacity as moderator of the perceptual disfluency effect
- P -hacking and Significance Stars
- Common misinterpretations of statistical significance and P-values in dairy research
- Using Bayes factor hypothesis testing in neuroscience to establish evidence of absence
- On not testing the foreign-language effect: A comment on Costa, Foucart, Arnon, Aparici, and Apesteguia (2014)
- Null hypothesis significance tests: A mix-up of two different theories, the basis for widespread confusion and numerous misinterpretations
- Bayesian two-interval test
- Infants' preference for speech is stable across the first year of life: Meta‐analytic evidence
- Interlimb coordination in Parkinson’s Disease is minimally affected by a visuospatial dual task
- The impact of using biased performance metrics on software defect prediction research
- Environmental Literature as Persuasion: An Experimental Test of the Effects of Reading Climate Fiction
- From Methodology to Practice
- Do VLMs Align Better with Humans than LLMs during Natural Reading?
- p-Hacking Inflates Type I Error Rates in the Error Statistical Approach but not in the Formal Inference Approach
- Why P Values Are Not a Useful Measure of Evidence in Statistical Significance Testing
- A Test by Any Other Name:PValues, Bayes Factors, and Statistical Inference
- p -Curve and Effect Size
- Connecting simple and precise P‐values to complex and ambiguous realities (includes rejoinder to comments on “Divergence vs. decision P‐values”)
- What’s in a Name: A Bayesian Hierarchical Analysis of the Name-Letter Effect
- The role of auditory and cognitive factors in understanding speech in noise by normal-hearing older listeners. [europepmc]
- Unifying treatments for depression: an application of the Free Energy Principle. [europepmc]
- Perils and pitfalls of reporting sex differences. [europepmc]
- Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. [europepmc]
- BNDF methylation in mothers and newborns is associated with maternal exposure to war trauma. [europepmc]
- The earth is flat ( p > 0.05): significance thresholds and the crisis of unreplicable research. [europepmc]
- Impact of Prefrontal Theta Burst Stimulation on Clinical Neuropsychological Tasks. [europepmc]
- Replication Study: Transcriptional amplification in tumor cells with elevated c-Myc. [europepmc]
- Large-scale replication study reveals a limit on probabilistic prediction in language comprehension. [europepmc]
- Using language input and lexical processing to predict vocabulary size. [europepmc]
- How do we assess a racial disparity in health? Distribution, interaction, and interpretation in epidemiological studies. [europepmc]
- Reconstructing the History of Polygenic Scores Using Coalescent Trees. [europepmc]
- The Longitudinal Association between Self-esteem and Depressive Symptoms in Adolescents: Separating between-person effects from within-person effects. [europepmc]
- Estimation for Better Inference in Neuroscience. [europepmc]
- Reporting and interpretation of results from clinical trials that did not claim a treatment difference: survey of four general medical journals. [europepmc]
- A large-scale genomic investigation of susceptibility to infection and its association with mental disorders in the Danish population. [europepmc]
- A synthetic dataset primer for the biobehavioural sciences to promote reproducibility and hypothesis generation. [europepmc]
- Noradrenergic-dependent functions are associated with age-related locus coeruleus signal intensity differences. [europepmc]
- Semantic and cognitive tools to aid statistical science: replace confidence and significance by compatibility and surprise. [europepmc]
- Inferential challenges when assessing racial/ethnic health disparities in environmental research. [europepmc]
- There is life beyond the statistical significance. [europepmc]
- Cortical alpha oscillations in cochlear implant users reflect subjective listening effort during speech-in-noise perception. [europepmc]
- Reporting and misreporting of sex differences in the biological sciences. [europepmc]
- Why and How to Account for Sex and Gender in Brain and Behavioral Research. [europepmc]
- Six persistent research misconceptions. [europepmc]
Discussions
Related