The Kappa Statistic in Reliability Studies: Use, Interpretation, and Sample Size Requirements
2005/03/01 by Julius Sim, Chris Wright, Chris C Wright · 37 citations
Decision Sciences · Mathematics · Medicine · #Reliability and Agreement in Measurement #Statistical Methods in Epidemiology #Inflammatory Biomarkers in Disease Prognosis
paper · pdf · doi:10.1093/ptj/85.3.257
Abstract
PURPOSE: This article examines and illustrates the use and interpretation of the kappa statistic in musculoskeletal research. SUMMARY OF KEY POINTS: The reliability of clinicians' ratings is an important consideration in areas such as diagnosis and the interpretation of examination findings. Often, these ratings lie on a nominal or an ordinal scale. For such data, the kappa coefficient is an appropriate measure of reliability. Kappa is defined, in both weighted and unweighted forms, and its use is illustrated with examples from musculoskeletal research. Factors that can influence the magnitude of kappa (prevalence, bias, and non-independent ratings) are discussed, and ways of evaluating the magnitude of an obtained kappa are considered. The issue of statistical testing of kappa is considered, including the use of confidence intervals, and appropriate sample sizes for reliability studies using kappa are tabulated. CONCLUSIONS: The article concludes with recommendations for the use and interpretation of kappa.
Citations
Cited by
- WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays
- Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework
- Relationship between body mass index percentile and skeletal maturation and dental development in orthodontic patients
- Analyzing Dataset Annotation Quality Management in the Wild
- The Hubble Image Similarity Project
- Hysteresis in streamflow‐water table relation provides a new classification system of rainfall‐runoff events
- Development of a research tool leveraging theoretical frameworks to better understand One Health systems thinking among livestock farmers
- Classification Model Evaluation Metrics
- Min-Mid-Max Scaling, Limits of Agreement, and Agreement Score
- Nomogram for sample size calculation on a straightforward basis for the kappa statistic
- Automated sleep stage identification system based on time–frequency analysis of a single EEG channel and random forest classifier
- Confident Privacy Decision-Making in IoT Environments
- Moral Minds in Gaming
- Inter-Rater Reliability of a 6-Item Movement Control Test Battery in Individuals With and Without Chronic Non-Specific Low Back Pain
- Case reports unlocked: Harnessing large language models to advance research on child maltreatment
- My Body, My Exoskeleton: Co-Designing Intersectional Visions of Robotic Augmentations through Drawings
- A comparative analysis of student, educator, and simulated parent ratings of video-recorded medical student consultations in pediatrics
- Ultra‐Processed Food Consumption, Mental Health, and Quality of Life in Adults: A Systematic Review Protocol
- Comparative Analysis of the Whistles of Three Oceanic Dolphins in the Comoros
- Optimized imaging prefiltering for enhanced image segmentation
- Architectural Degradation: Definition, Motivations, Measurement and Remediation Approaches
- Validity and Reliability of Palpatory Clinical Tests of Sacroiliac Joint Mobility: A Systematic Review and Meta-analysis
- Improving Struggling Fifth-Grade Students’ Understanding of Fractions: A Randomized Controlled Trial of an Intervention That Stresses Both Concepts and Procedures
- Susceptibility of Citrus germplasm to canker caused by Neofusicoccum parvum
- Group discussions improve reliability and validity of rated categories based on qualitative data from systematic review
- Meta-Fair: AI-Assisted Fairness Testing of Large Language Models
- Applying Science Models for Search
- Decide less, communicate more: On the construct validity of end-to-end fact-checking in medicine
- CCISolver: End-to-End Detection and Repair of Method-Level Code-Comment Inconsistency
- MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
- Learning to Diversify via Weighted Kernels for Classifier Ensemble
- Identifying Offensive Expressions of Opinion in Context
- Cataloguing Hugging Face Models to Software Engineering Activities: Automation and Findings
- Reliability of physical examination tests used in the assessment of patients with shoulder problems: a systematic review
- The use of the MK5 Mobility Classes to improve safe patient handling: a reliability study
- Intraexaminer and Interexaminer Reliability of Manual Palpation and Pressure Algometry of the Lower Limb Nerves in Asymptomatic Subjects
- Inter(sectional) Alia(s): Ambiguity in Voice Agent Identity via Intersectional Japanese Self-Referents
Related