2003/08/01 by Raymond Hubbard, M. J. Bayarri · 1 citation
Mathematics · Physics and Astronomy · Psychology · #Probability and Statistical Research #Advanced Statistical Methods and Models #Statistical Mechanics and Entropy #Statistical hypothesis testing #Statistical inference #p-value #Confusion #Inductive reasoning #Mathematics #Statistics #Type I and type II errors #Statistical significance #Inference #Econometrics #Interpretation (philosophy) #Mathematical economics #Epistemology #Psychology #Philosophy #Linguistics
paper · doi:10.1198/0003130031856
openalex publication_date 2003/08/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/29
Confusion surrounding the reporting and interpretation of results of classical statistical tests is widespread among applied researchers, most of whom erroneously believe that such tests are prescribed by a single coherent theory of statistical inference. This is not the case: Classical statistical testing is an anonymous hybrid of the competing and frequently contradictory approaches formulated by R. A. Fisher on the one hand, and Jerzy Neyman and Egon Pearson on the other. In particular, there is a widespread failure to appreciate the incompatibility of Fisher's evidential p value with the Type I error rate, α, of Neyman-Pearson statistical orthodoxy. The distinction between evidence (p's) and error (α's) is not trivial. Instead, it reflects the fundamental differences between Fisher's ideas on significance testing and inductive inference, and Neyman-Pearson's views on hypothesis testing and inductive behavior. The emphasis of the article is to expose this incompatibility, but we also briefly note a possible reconciliation.