Note on the Sampling Error of the Difference Between Correlated Proportions or Percentages
1947/06/01 by Quinn McNemar · 123 citations
Mathematics · #Advanced Statistical Methods and Models
paper · doi:10.1007/bf02295996
Abstract
Two formulas are presented for judging the significance of the difference between correlated proportions. The chi square equivalent of one of the developed formulas is pointed out.
Cited by
- Asymptotic equivalence of paired Hotelling test and conditional logistic\n regression
- Modeling the zebrafish gut microbiome’s resistance and sensitivity to climate change and parasite infection
- Quantize with Confidence? An Empirical Study of Quantization for Code Generation
- Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?
- From Bit-Position Sensitivity to Unequal Error Protection for DNN Inference Memory
- In-Context Learning for Wound Classification with Small Multimodal Language Models
- When Directional Accuracy Lies: A Base-Rate-Honest Benchmark for LoRA-Adapted TimesFM on Equity Forecasting
- Style over Substance: A Shortcut Audit of Emotion-Description Preference Evaluation
- Grounded verification of chemical and materials reasoning: detection is the bottleneck
- When Model Merging Rivals Joint Multi-Task Reinforcement Learning: A Task-Vector Geometry Analysis
- Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
- A Transportable Threshold-Based Framework for Interpretable Classification of Medical Data
- Reading Between the Dots: Decoding Hidden Computation across Filler Tokens
- Surface Fairness, Deep Bias: A Comparative Study of Bias in Language Models
- Logical Concepts of (Im)possibility Guide Young Children's Decision‐Making
- Brain functional connectivity, but not neuroanatomy, captures the interrelationship between sex and gender in preadolescents
- Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models
- Beyond the Tip of the Iceberg: Assessing Coherence of Text Classifiers
- Robust Cross-lingual Hypernymy Detection using Dependency Context
- Replicability Analysis for Natural Language Processing: Testing Significance with Multiple Datasets
- Unsupervised Learning of Parsimonious General-Purpose Embeddings for User and Location Modelling
- Harbor Porpoise and Beluga Whale Habitat Use in the Saguenay‐St. Lawrence Marine Park (Canada) Revealed by a Combination of Visual and Acoustic Survey
- Calibrated Tree-Neural Fusion for Fine-Grained Vegetation Community Classification
- KBF: Knowledge Boundary as Fingerprint for Language Model and Black-Box API Auditing
- Algorithm for Interpretable Graph Features via Motivic Persistent Cohomology
- Identifying Features Associated with Bias Against 93 Stigmatized Groups in Language Models and Guardrail Model Safety Mitigation
- Toxicity Ahead: Forecasting Conversational Derailment on GitHub
- Imitation Game: Reproducing Deep Learning Bugs Leveraging an Intelligent Agent
- Prompt Repetition Improves Non-Reasoning LLMs
- IaC Generation with LLMs: An Error Taxonomy and A Study on Configuration Knowledge Injection
- Verifying Rumors via Stance-Aware Structural Modeling
- Error-Driven Prompt Optimization for Arithmetic Reasoning
- FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Reformulate, Retrieve, Localize: Agents for Repository-Level Bug Localization
- Inferring laminar origins of MEG signals with optically pumped magnetometers (OPMs): A simulation study
- Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
- Multi-Agent Code Verification via Information Theory
- Detecting Viruses in Contact Networks with Unreliable Detectors
- Convolutional Neural Networks for Large-Scale Remote-Sensing Image Classification
- AutoMalDesc: Large-Scale Script Analysis for Cyber Threat Research
- Multi-task Domain Adaptation for Sequence Tagging
- ExPairT-LLM: Exact Learning for LLM Code Selection by Pairwise Queries
- An Exploratory Eye Tracking Study on How Developers Classify and Debug Python Code in Different Paradigms
- On Stealing Graph Neural Network Models
- Testing the Testers: Human-Driven Quality Assessment of Voice AI Testing Platforms
- Robust Alignment of the Human Embryo in 3D Ultrasound using PCA and an Ensemble of Heuristic, Atlas-based and Learning-based Classifiers Evaluated on the Rotterdam Periconceptional Cohort
- A Large-Scale Comprehensive Measurement of AI-Generated Code in Real-World Repositories
- ECONET: Effective Continual Pretraining of Language Models for Event Temporal Reasoning
- Classifier comparison using precision
- Try Again, Don't Look Back: Blind Resampling Outperforms Self-Repair in Small Code Models
- Short text classification with machine learning in the social sciences: The case of climate change on Twitter
- Invisible Stripes? A Field Experiment on the Disclosure of a Criminal Record in the British Labour Market and the Potential Effects of Introducing Ban-The-Box Policies
- Changes in growth of the Naru eagle ray Aetobatus narutobiei in Ariake Bay, based on over two decades of monitoring under fishing pressure
- Identifying Causal Genotype–Phenotype Relationships for Population‐Sampled Parent–Child Trios
- Forensic Reproducibility Audit of a Radiology Vision-Language Model Benchmark: From Intended Protocol to Released Artifact
- Biases in the Blind Spot: Detecting What LLMs Fail to Mention
- Adaptive EEG-based stroke diagnosis with a GRU-TCN classifier and deep Q-learning thresholding
- Between dunes and estuary: Forecasting mangrove forest change on primate culture and isolated livelihoods in Maranhão, Brazil
- On the Faithfulness of Visual Thinking: Measurement and Enhancement
- Building Trust in Clinical LLMs: Bias Analysis and Dataset Transparency
- Combining Distantly Supervised Models with In Context Learning for Monolingual and Cross-Lingual Relation Extraction
- When Old Meets New: Evaluating the Impact of Regression Tests on SWE Issue Resolution
- Inaccessibility-Inside Theorem for Point in Polygon
- GADA: Graph Attention-based Detection Aggregation for Ultrasound Video Classification
- RefFilter: Improving Semantic Conflict Detection via Refactoring-Aware Static Analysis
- To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Re-Identifying Kākā with AI-Automated Video Key Frame Extraction
- Drawing Conclusions from Draws: Rethinking Preference Semantics in Arena-Style LLM Evaluation
- Reasoning under Vision: Understanding Visual-Spatial Cognition in Vision-Language Models for CAPTCHA
- Reinforcement Learning for Clinical Reasoning: Aligning LLMs with ACR Imaging Appropriateness Criteria
- Voice Evaluation of Reasoning Ability: Diagnosing the Modality-Induced Performance Gap
- BloomAPR: A Bloom's Taxonomy-based Framework for Assessing the Capabilities of LLM-Powered APR Solutions
- Non-Invasive Detection of PROState Cancer with Novel Time-Dependent Diffusion MRI and AI-Enhanced Quantitative Radiological Interpretation: PROS-TD-AI
- Uncertainty-aware visualization and integration for maritime route safety assessment
- Gender Equality Plans and Inclusiveness in the European Research Area
- ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents
- Comparison of whole blood on filter strips with serum for avian influenza virus antibody detection in wild birds
- Knowledge distillation through geometry-aware representational alignment
- Urban land use and land cover classification with interpretable machine learning – A case study using Sentinel-2 and auxiliary data
- A Bayesian Approach to Age Estimation in Modern Americans from the Clavicle*
- Quantifying the Impact of Structured Output Format on Large Language Models through Causal Inference
- In AI Sweet Harmony: Sociopragmatic Guardrail Bypasses and Evaluation-Awareness in OpenAI gpt-oss-20b
- Benchmarking LLM Competence on Logical Inference over Probability Operators
- The State of Peptide Detectability in Computational Proteomics and Guidelines for AI Applications
- PLM-interact: extending protein language models to predict protein-protein interactions
- Evaluating LLM Agents on Automated Software Analysis Tasks
- Relational Scene Graphs for Object Grounding of Natural Language Commands
- DL-QC-fNIRS: a deep learning tool for automated quality control in functional near-infrared spectroscopy signals
- Assessing model-driven mutation testing of Java bytecode
- Machine learning-based ensemble species distribution models to guide monitoring and survey design for offshore wind
- Correlation or Causation: Analyzing the Causal Structures of LLM and LRM Reasoning Process
- GnnXemplar: Exemplars to Explanations -- Natural Language Rules for Global GNN Interpretability
- Robustness of Neurosymbolic Reasoners on First-Order Logic Problems
- Towards Optimal Convolutional Transfer Learning Architectures for Breast Lesion Classification and ACL Tear Detection
- Does medical school cause depression or do medical students already begin their studies depressed? A longitudinal study over the first semester about depression and influencing factors
- Improving patch-based scene text script identification with ensembles of\n conjoined networks
- Multiple imputation: a primer
- An Alternative Measure of Effect Size for Cochran's <i>Q</i> Test for Related Proportions
- Acquiescence Bias in Large Language Models
- What Were You Thinking? An LLM-Driven Large-Scale Study of Refactoring Motivations in Open-Source Projects
- Scam2Prompt: A Scalable Framework for Auditing Malicious Scam Endpoints in Production LLMs
- Signs of Struggle: Spotting Cognitive Distortions across Language and Register
- Can News Predict the Direction of Oil Price Volatility? A Language Model Approach with SHAP Explanations
- Identifying and Answering Questions with False Assumptions: An Interpretable Approach
- A Generic Framework for Assessing the Performance Bounds of Image\n Feature Detectors
- The Hidden Cost of Readability: How Code Formatting Silently Consumes Your LLM Budget
- Is GPT-OSS Good? A Comprehensive Evaluation of OpenAI's Latest Open Source Models
- Cost-Aware Contrastive Routing for LLMs
- Design and Validation of a Responsible Artificial Intelligence-based System for the Referral of Diabetic Retinopathy Patients
- Lameness detection in dairy cows using pose estimation and bidirectional LSTMs
- Over-Squashing in GNNs and Causal Inference of Rewiring Strategies
- TEN: Table Explicitization, Neurosymbolically
- Evaluating Large Language Models as Expert Annotators
- Exploring Causal Effect of Social Bias on Faithfulness Hallucinations in Large Language Models
- Heterogeneous Prompting and Execution Feedback for SWE Issue Test Generation and Selection
- Representation Learning for Words and Entities
- Artificial Intelligence-Based Classification of Spitz Tumors
- Empirical Evaluation of AI-Assisted Software Package Selection: A Knowledge Graph Approach
- Improving Generalization in Coreference Resolution via Adversarial Training
- Building and Aligning Comparable Corpora
- LiveMCPBench: Can Agents Navigate an Ocean of MCP Tools?
- Argumentatively Coherent Judgmental Forecasting
- Evaluating the added predictive ability of a new marker: From area under the ROC curve to reclassification and beyond
Related