Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages
1947/06/01 by Quinn McNemar · 4,394 citations
Mathematics · #Advanced Statistical Methods and Models #Mathematics #Sampling (signal processing) #Statistical significance #Statistics
paper · doi:10.1007/bf02295996
published in Psychometrika 12(2), 153-157 (Springer Science+Business Media)
openalex publication_date 1947/06/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/06
Abstract
Two formulas are presented for judging the significance of the difference between correlated proportions. The chi square equivalent of one of the developed formulas is pointed out.
Cited by
- Asymptotic equivalence of paired Hotelling test and conditional logistic\n regression
- Modeling the zebrafish gut microbiome’s resistance and sensitivity to climate change and parasite infection
- Quantize with Confidence? An Empirical Study of Quantization for Code Generation
- Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting
- Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation
- Grounded verification of chemical and materials reasoning: detection is the bottleneck
- When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
- Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
- A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
- Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
- Logical Concepts of (Im)possibility Guide Young Children's Decision‐Making
- Brain functional connectivity, but not neuroanatomy, captures the interrelationship between sex and gender in preadolescents
- Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
- Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
- Robust Cross-lingual Hypernymy Detection using Dependency Context
- Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
- Unsupervised Learning of Parsimonious General-Purpose Embeddings for User and Location Modelling
- Harbor Porpoise and Beluga Whale Habitat Use in the Saguenay‐St. Lawrence Marine Park (Canada) Revealed by a Combination of Visual and Acoustic Survey
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
- Algorithm for Interpretable Graph Features via Motivic Persistent Cohomology
- Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
- Toxicity Ahead: Forecasting Conversational Derailment on GitHub
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
- Prompt Repetition Improves Non-Reasoning LLMs
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Verifying Rumors via Stance-Aware Structural Modeling
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization
- Inferring laminar origins of MEG signals with optically pumped magnetometers (OPMs): A simulation study
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
- Multi-Agent Code Verification via Information Theory
- Detecting Viruses in Contact Networks with Unreliable Detectors
- Convolutional Neural Networks for Large-Scale Remote-Sensing Image Classification
- AutoMalDesc: Large-Scale Script Analysis for Cyber Threat Research
- Multi-task Domain Adaptation for Sequence Tagging
- ExPairT-LLM: Exact Learning for LLM Code Selection by Pairwise Queries
- An Exploratory Eye Tracking Study on How Developers Classify and Debug Python Code in Different Paradigms
- On Stealing Graph Neural Network Models
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- Robust Alignment of the Human Embryo in 3D Ultrasound using PCA and an Ensemble of Heuristic, Atlas-based and Learning-based Classifiers Evaluated on the Rotterdam Periconceptional Cohort
- A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning
- Classifier comparison using precision
- Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
- Short text classification with machine learning in the social sciences: The case of climate change on Twitter
- Invisible Stripes? A Field Experiment on the Disclosure of a Criminal Record in the British Labour Market and the Potential Effects of Introducing Ban-The-Box Policies
- Changes in growth of the Naru eagle ray Aetobatus narutobiei in Ariake Bay, based on over two decades of monitoring under fishing pressure
- Identifying Causal Genotype–Phenotype Relationships for Population‐Sampled Parent–Child Trios
- Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- Adaptive EEG-based stroke diagnosis with a GRU-TCN classifier and deep Q-learning thresholding
- Between dunes and estuary: Forecasting mangrove forest change on primate culture and isolated livelihoods in Maranhão, Brazil
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
- Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction
- When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution
- Inaccessibility-Inside Theorem for Point in Polygon
- GADA: Graph Attention-based Detection Aggregation for Ultrasound Video Classification
- RefFilter: Improving Semantic Conflict Detection via Refactoring-Aware Static Analysis
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
- Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
- Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
- Reinforcement Learning for Clinical Reasoning: Aligning LLMs with ACR Imaging Appropriateness Criteria
- Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
- BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
- Non-Invasive Detection of PROState Cancer with Novel Time-Dependent Diffusion MRI and AI-Enhanced Quantitative Radiological Interpretation: PROS-TD-AI
- Uncertainty-aware visualization and integration for maritime route safety assessment
- Gender Equality Plans and Inclusiveness in the European Research Area
- ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents
- Comparison of whole blood on filter strips with serum for avian influenza virus antibody detection in wild birds
- Knowledge distillation through geometry-aware representational alignment
- Urban land use and land cover classification with interpretable machine learning – A case study using Sentinel-2 and auxiliary data
- A Bayesian Approach to Age Estimation in Modern Americans from the Clavicle*
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
- Benchmarking LLM Competence on Logical Inference over Probability Operators
- The State of Peptide Detectability in Computational Proteomics and Guidelines for AI Applications
- PLM-interact: extending protein language models to predict protein-protein interactions
- Evaluating LLM Agents on Automated Software Analysis Tasks
- Relational Scene Graphs for Object Grounding of Natural Language Commands
- DL-QC-fNIRS: a deep learning tool for automated quality control in functional near-infrared spectroscopy signals
- Assessing model-driven mutation testing of Java bytecode
- Machine learning-based ensemble species distribution models to guide monitoring and survey design for offshore wind
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- GnnXemplar: Exemplars to Explanations -- Natural Language Rules for Global GNN Interpretability
- Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- Towards Optimal Convolutional Transfer Learning Architectures for Breast Lesion Classification and ACL Tear Detection
- Does medical school cause depression or do medical students already begin their studies depressed? A longitudinal study over the first semester about depression and influencing factors
- Improving patch-based scene text script identification with ensembles of conjoined networks
- Multiple imputation: a primer
- An Alternative Measure of Effect Size for Cochran's Q Test for Related Proportions
- Acquiescence Bias in Large Language Models
- What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
- Signs of Struggle: Spotting Cognitive Distortions across Language and Register
- Can News Predict the Direction of Oil Price Volatility? A Language Model Approach with SHAP Explanations
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- A Generic Framework for Assessing the Performance Bounds of Image Feature Detectors
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Cost-Aware Contrastive Routing for LLMs
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- Lameness detection in dairy cows using pose estimation and bidirectional LSTMs
- Over-Squashing in GNNs and Causal Inference of Rewiring Strategies
- TEN: Table Explicitization, Neurosymbolically
- Evaluating Large Language Models as Expert Annotators
- Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
- Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
- Representation Learning for Words and Entities
- Artificial Intelligence-Based Classification of Spitz Tumors
- Empirical Evaluation of AI-Assisted Software Package Selection: A Knowledge Graph Approach
- Improving Generalization in Coreference Resolution via Adversarial Training
- Building and Aligning Comparable Corpora
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- Argumentatively Coherent Judgmental Forecasting
- Evaluating the added predictive ability of a new marker: From area under the ROC curve to reclassification and beyond
- On Arbitrary Predictions from Equally Valid Models
- MemoCoder: Automated Function Synthesis using LLM-Supported Agents
- Impact of Geant4's Electromagnetic Physics Constructors on Accuracy and Performance of Simulations for Rare Event Searches
- Pathology-Aware Generative Adversarial Networks for Medical Image Augmentation
- AEGIS: A Backup Reflex for Physical AI
- MelT: A Portable, Single-GEMM Mel Audio Frontend via Non-Uniform DFT with Measured Latency and Energy Gains on GPUs
- MONITRS: Multimodal Observations of Natural Incidents Through Remote Sensing
- Written Justifications are Key to Aggregate Crowdsourced Forecasts
- Detecting Asks in SE attacks: Impact of Linguistic and Structural Knowledge
- Comparing Attention-based Convolutional and Recurrent Neural Networks:\n Success and Limitations in Machine Reading Comprehension
- Occam Factor for Gaussian Models With Unknown Variance Structure
- Preventing Premature Commitment in Coding Agents with an Evidence-Conditioned Execution Layer
- CARA: Exact Local Repair with Fresh One-Action Certification for Cloud Consolidation
- Bridging the Plausibility-Validity Gap by Fine-Tuning a Reasoning-Enhanced LLM for Chemical Synthesis and Discovery
- IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care
- Log-Likelihood Ratio Minimizing Flows: Towards Robust and Quantifiable Neural Distribution Alignment
- Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
- Cross-lingual Models of Word Embeddings: An Empirical Comparison
- Think Before You Code: Dual Reasoning for the NLSafety-Utility Trade-Off in LLM Code Generation
- Exploring a Hybrid Deep Learning Approach for Anomaly Detection in Mental Healthcare Provider Billing: Addressing Label Scarcity through Semi-Supervised Anomaly Detection
- Mathematics Isn't Culture-Free: Probing Cultural Gaps via Entity and Scenario Perturbations
- Hospital-related healthcare expenditure of impending versus completed pathological femur fractures: a propensity score matched study of 265 patients
- Effect of artificial intelligence-aided differentiation of adenomatous and non-adenomatous colorectal polyps at CT colonography on radiologists’ therapy management
- Lemmatization as a Classification Task: Results from Arabic across Multiple Genres
- CopulaSMOTE: A Copula-Based Oversampling Approach for Imbalanced Classification in Diabetes Prediction
- Classification of Multi-Parametric Body MRI Series Using Deep Learning
- LLM vs. SAST: A Technical Analysis on Detecting Coding Bugs of GPT4-Advanced Data Analysis
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- JavaNPST: Nonparametric Statistical Tests in Java
- A Hierarchy of Limitations in Machine Learning
- Givenness Hierarchy Theoretic Cognitive Status Filtering
- Students under lockdown: Comparisons of students’ social networks and mental health before and during the COVID-19 crisis in Switzerland
- Characterization of Posidonia Oceanica Seagrass Aerenchyma through Whole Slide Imaging: A Pilot Study
- ESTER: A Machine Reading Comprehension Dataset for Event Semantic Relation Reasoning
- A Copula Based Supervised Filter for Feature Selection in Diabetes Risk Prediction Using Machine Learning
- Graph Classification using Signal-Subgraphs: Applications in Statistical Connectomics
- MultiPhishGuard: An LLM-based Multi-Agent System for Phishing Email Detection
- FairMedQA: Benchmarking Bias in Large Language Models for Medical Question Answering
- Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models
- Can a domain-specific language improve program structure comprehension of data pipelines? A mixed-methods study
- HALT: Verification-Aware Stopping for Retrieval-Augmented Search Agents
- Cross-validation Confidence Intervals for Test Error
- Cross-Fitted Residual Utility for Primary-Preserving Cognitive Decision Correction in Automatic Modulation Classification
- HPFA: Hypergraph-Based Paired Failure Attribution for LLM Reasoning
- Evaluation of a Trauma‐Focused Group Intervention for Unaccompanied Young Refugees: A Pilot Study
- Doc2CI: A Multi-Service Study of CI Configuration Generation Using Large Language Models
- Post-mega-event domestic tourism development: Attitudinal transformation and market segmentation among young Qatari women following the 2022 FIFA World Cup
- Recurrent Connectivity Aids Recognition of Partly Occluded Objects
- Does It Capture STEL? A Modular, Similarity-based Linguistic Style Evaluation Framework
- SPARC-Rad: A Multimodal Benchmark Dataset and Evaluation Pipeline for Spatial and Anatomical Reasoning in Radiology Vision-Language Models
- Multi-Plane Vision Transformer for Hemorrhage Classification Using Axial and Sagittal MRI Data
- Towards Explainable Fact Checking
- DBLFace: Domain-Based Labels for NIR-VIS Heterogeneous Face Recognition
- Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
- From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents
- Intrinsic tests for the equality of two correlated proportions
- DeformSyncNet: Deformation Transfer via Synchronized Shape Deformation Spaces
- Reading Proficiency in Elementary: Considering Statewide Testing, Teacher Ratings and Rankings, and Reading Curriculum-Based Measurement
- Single Reading with Computer-Aided Detection for Screening Mammography
- Reasoning Capabilities and Invariability of Large Language Models
- Discourse-aware rumour stance classification in social media using sequential classifiers
- Can singing rate be used to predict male breeding status of forest songbirds? A comparison of three calibration models
- BioTool: A Comprehensive Tool-Calling Dataset for Enhancing Biomedical Capabilities of Large Language Models
- Same Meaning, Different Scores: Lexical and Syntactic Sensitivity in LLM Evaluation
- Improving Code Generation via Small Language Model-as-a-judge
- PlotChain: Deterministic Checkpointed Evaluation of Multimodal LLMs on Engineering Plot Reading
- ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers
- Did you miss it? Automatic lung nodule detection combined with gaze information improves radiologists' screening performance
- Human Values in a Single Sentence: Moral Presence, Hierarchies, and Transformer Ensembles on the Schwartz Continuum
- Native Language Identification using Stacked Generalization
- Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
- Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models
- Surrogate Substitution Preserves PHI Detectability: A Multi-Detector Equivalence Study
- TRACE: Textual Relevance Augmentation and Contextual Encoding for Multimodal Hate Detection
- Syllabic quantity patterns as rhythmic features for Latin authorship attribution
- A survey on concept drift adaptation
- Learning under Concept Drift: A Review
- Scrouting: Cost-Aware Routing of Coding Agents by Scouting the Repository First
- An Approach for Embedding-Guided Function Reuse Detection in Embedded C Software
- LookAhead: Augmenting Crowdsourced Website Reputation Systems With Predictive Modeling
- Unsupervised transfer learning for anomaly detection: Application to complementary operating condition transfer
- Significativity Indices for Agreement Values
- Tratto: A Neuro-Symbolic Approach to Deriving Axiomatic Test Oracles
- Towards Competence-Based Management for Open Source Software Projects
- Enhancing accuracy and privacy in speech-based depression detection through speaker disentanglement
- Assessing how hyperparameters impact Large Language Models' sarcasm detection performance
- Large deformation diffeomorphism and momentum based hippocampal shape discrimination in dementia of the Alzheimer type. [europepmc]
- Top-down control of human visual cortex by frontal and parietal cortex in anticipatory visual spatial attention. [europepmc]
- Perilesional brain oedema and seizure activity in patients with calcified neurocysticercosis: a prospective cohort and nested case-control study. [europepmc]
- Surgical versus nonoperative treatment for lumbar disc herniation: four-year results for the Spine Patient Outcomes Research Trial (SPORT). [europepmc]
- Prediction of 6-month survival of nursing home residents with advanced dementia using ADEPT vs hospice eligibility guidelines. [europepmc]
- The social ecological model as a framework for determinants of 2009 H1N1 influenza vaccine uptake in the United States. [europepmc]
- Prefrontal-occipitoparietal coupling underlies late latency human neuronal responses to emotion. [europepmc]
- Intravenous contrast material-induced nephropathy: causal or coincident phenomenon? [europepmc]
- High HIV testing uptake and linkage to care in a novel program of home-based HIV counseling and testing with facilitated referral in KwaZulu-Natal, South Africa. [europepmc]
- Evaluation of urovysion and cytology for bladder cancer detection: a study of 1835 paired urine samples with clinical and histologic correlation. [europepmc]
- The McNemar test for binary matched-pairs data: mid-p and asymptotic are better than exact conditional. [europepmc]
- Host gene expression classifiers diagnose acute respiratory illness etiology. [europepmc]
- Predicting Malignant Nodules from Screening CT Scans. [europepmc]
- Recurrent Convolutional Neural Networks: A Better Model of Biological Object Recognition. [europepmc]
- iBCE-EL: A New Ensemble Learning Framework for Improved Linear B-Cell Epitope Prediction. [europepmc]
- Meta-4mCpred: A Sequence-Based Meta-Predictor for Accurate DNA 4mC Site Prediction Using Effective Feature Representation. [europepmc]
- Real-time decoding of question-and-answer speech dialogue using human cortical activity. [europepmc]
- Deep neural networks and kernel regression achieve comparable accuracies for functional connectivity prediction of behavior and demographics. [europepmc]
- Automatic diagnosis of the 12-lead ECG using a deep neural network. [europepmc]
- Criteria for defining interictal epileptiform discharges in EEG: A clinical validation study. [europepmc]
- Multiparametric MRI for Prostate Cancer Characterization: Combined Use of Radiomics Model with PI-RADS and Clinical Parameters. [europepmc]
- Students under lockdown: Comparisons of students' social networks and mental health before and during the COVID-19 crisis in Switzerland. [europepmc]
- A standardized framework for testing the performance of sleep-tracking technology: step-by-step guidelines and open-source code. [europepmc]
- TSEBRA: transcript selector for BRAKER. [europepmc]
- A foundation model for clinical-grade computational pathology and rare cancers detection. [europepmc]
Related