The Equivalence of Weighted Kappa and the Intraclass Correlation Coefficient as Measures of Reliability
1973/10/01 by Joseph L. Fleiss, Jacob Cohen · 3,373 citations
Decision Sciences · Mathematics · Medicine · Psychology · #Advanced Statistical Methods and Models #Cohen's kappa #Correlation #Correlation coefficient #Correlation ratio #Discrete mathematics #Econometrics #Equivalence (formal languages) #Hemodynamic Monitoring and Therapy #Interclass correlation #Intraclass correlation #Kappa #Mathematics #Physics #Psychology #Psychometrics #Reliability (semiconductor) #Reliability and Agreement in Measurement #Statistics
paper · doi:10.1177/001316447303300309
published in Educational and Psychological Measurement 33(3), 613-619 (SAGE Publishing)
openalex publication_date 1973/10/01 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05
Citations
Cited by
- The reliability of the functional independence measure: A quantitative review
- A Note on the Interpretation of Weighted Kappa and its Relations to Other Rater Agreement Statistics for Metric Scales
- The Boston bowel preparation scale: a valid and reliable instrument for colonoscopy-oriented research
- The Measurement of Observer Agreement for Categorical Data
- The Manual Ability Classification System (MACS) for children with cerebral palsy: scale development and evidence of validity and reliability
- A Revised Diagnostic Classification of Canine Glioma: Towards Validation of the Canine Glioma Patient as a Naturally Occurring Preclinical Model for Human Glioma
- Precision of Health-Related Quality-of-Life Data Compared With Other Clinical Measures
- Dynamic Commonsense Coordination for Empathetic Response Generation
- Inter-Rater Reliability Methods in Qualitative Case Study Research
- 3D multi‐view squeeze‐and‐excitation convolutional neural network for lung nodule classification
- Controllability of Stressful Events and Satisfaction With Spouse Support Behaviors
- Using Natural Language Processing to Automatically Detect Self-Admitted Technical Debt
- Textverständlichkeit und kognitive Belastung beim Lernen mit Text und Hypertext
- SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
- Enhancing Job Matching: Occupation, Skill and Qualification Linking with the ESCO and EQF taxonomies
- Cross-replication Reliability -- An Empirical Approach to Interpreting Inter-rater Reliability
- Fantastic Bugs and Where to Find Them in AI Benchmarks
- De-identification of Privacy-related Entities in Job Postings
- Learning Neural Templates for Recommender Dialogue System
- Bien Educado: Measuring the social behaviors of Mexican American children
- MultiOpEd: A Corpus of Multi-Perspective News Editorials
- Consultorio Médico Local Tipo I. Vícar, Almería
- The Paris 1976 Wine Tastings Revisited Once More: Comparing Ratings of Consistent and Inconsistent Tasters
- Identifying & Interactively Refining Ambiguous User Goals for Data Visualization Code Generation
- CFVBench: A Comprehensive Video Benchmark for Fine-grained Multimodal Retrieval-Augmented Generation
- Whose Opinions Matter? Perspective-aware Models to Identify Opinions of Hate Speech Victims in Abusive Language Detection
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- A Note on the Use of Categorical Subscores
- QUARTZ : QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization
- Beyond kappa: A review of interrater agreement measures
- AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
- Impact of schematic representations of road maps on pedestrian route planning
- Psychometric properties of the EQ-5D-5L: a systematic review of the literature
- Testing the WHO Hand Hygiene Self-Assessment Framework for usability and reliability
- Resilients, Overcontrollers, and Undercontrollers: The replicability of the three personality prototypes across informants
- T2R-bench: A Benchmark for Generating Article-Level Reports from Real World Industrial Tables
- Self-Regulation Profiles and Social Competence in Early Childhood: A Person-Centered Approach
- Generating Multiple Diverse Responses with Multi-Mapping and Posterior Mapping Selection
- Discovering the Potential of Automated Phraseological Interference Error Detection: A Transformer-Based Approach
- Progressive Open-Domain Response Generation with Multiple Controllable Attributes
- SentiPers: A Sentiment Analysis Corpus for Persian
- A Multi-Turn Emotionally Engaging Dialog Model
- Guiding Variational Response Generator to Exploit Persona
- GitHub Discussions: An Exploratory Study of Early Adoption
- From Detection of Individual Metastases to Classification of Lymph Node Status at the Patient Level: The CAMELYON17 Challenge
- SkillSpan: Hard and Soft Skill Extraction from English Job Postings
- Intraclass Correlation Coefficient (ICC): A Framework for Monitoring and Assessing Performance of Trained Sensory Panels and Panelists
- Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents
- The Thin Line Between Comprehension and Persuasion in LLMs
- PLATO: Pre-trained Dialogue Generation Model with Discrete Latent Variable
- The International Caries Detection and Assessment System (ICDAS): an integrated system for measuring dental caries
- Domain Adaptative Causality Encoder
- Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
- Seeing Things from a Different Angle: Discovering Diverse Perspectives about Claims
- Cue-word Driven Neural Response Generation with a Shrinking Vocabulary
- EmailSum: Abstractive Email Thread Summarization
- Learning Causal Bayesian Networks from Text
- ProphetNet-X: Large-Scale Pre-training Models for English, Chinese, Multi-lingual, Dialog, and Code Generation
- DeepDialogue: A Multi-Turn Emotionally-Rich Spoken Dialogue Dataset
- Guidelines, criteria, and rules of thumb for evaluating normed and standardized assessment instruments in psychology.
- How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control
- Reliability of the Goutallier Classification in Quantifying Muscle Fatty Degeneration in the Lumbar Multifidus Using Magnetic Resonance Imaging
- DecoupledESC: Enhancing Emotional Support Generation via Strategy-Response Decoupled Preference Optimization
- Improving Medical Image Classification with Label Noise Using Dual-uncertainty Estimation
- CDL: Curriculum Dual Learning for Emotion-Controllable Response Generation
- Behavior of agreement measures in the presence of zero cells and biased marginal distributions
- “Island-shape” Fractures of Lister’s tubercle have an increased risk of delayed extensor pollicis longus rupture in distal radial fractures
- Maria: A Visual Experience Powered Conversational Agent
- A Re-analysis of the Reliability of Psychiatric Diagnosis
- Reliability of Scores on the Stroke Rehabilitation Assessment of Movement (STREAM) Measure
- Can LLMs Accurately Score Medical Diagnoses and Clinical Reasoning?
- Cuba: Exploring the History of Admixture and the Genetic Basis of Pigmentation Using Autosomal and Uniparental Markers
- An Examination of Interrater Reliability for Scoring the Rorschach Comprehensive System in Eight Data Sets
- The Effect of Number of Rating Scale Categories on Levels of Interrater Reliability : A Monte Carlo Investigation
- The role of teaching presence in students’ behavioral engagement
- On building an automated responding system for app reviews: What are the characteristics of reviews and their responses?
- The many dimensions of truthfulness: Crowdsourcing misinformation assessments on a multidimensional scale
- Resurrecting Socrates in the Age of AI: A Study Protocol for Evaluating a Socratic Tutor to Support Research Question Development in Higher Education
- Psychometric characteristics of the Spanish version of instruments to measure neck pain disability. [europepmc]
- Reliability of Ashworth and Modified Ashworth scales in children with spastic cerebral palsy. [europepmc]
- Agreement, reliability and validity in 3 shoulder questionnaires in patients with rotator cuff disease. [europepmc]
- Effect of rater training on reliability and accuracy of mini-CEX scores: a randomized, controlled trial. [europepmc]
- Validity/reliability of PHQ-9 and PHQ-2 depression scales among adults living with HIV/AIDS in western Kenya. [europepmc]
- Reliability and validity of triage systems in paediatric emergency care. [europepmc]
- Feasibility, reliability, and validity of the EQ-5D-Y: results from a multinational study. [europepmc]
- Health care providers underestimate symptom intensities of cancer patients: a multicenter European study. [europepmc]
- Prevention of ventilator-associated pneumonia, mortality and all intensive care unit acquired infections by topically applied antimicrobial or antiseptic agents: a meta-analysis of randomized controlled trials in intensive care units. [europepmc]
- Texture analysis of cartilage T2 maps: individuals with risk factors for OA have higher and more heterogeneous knee cartilage MR T2 compared to normal controls--data from the osteoarthritis initiative. [europepmc]
- Shear wave elastography for breast masses is highly reproducible. [europepmc]
- Enhancing medical students' communication skills: development and evaluation of an undergraduate training program. [europepmc]
- Percutaneous and surgical tracheostomy in critically ill adult patients: a meta-analysis. [europepmc]
- Tracking of overweight and obesity from early childhood to adolescence in a population-based cohort - the Tromsø Study, Fit Futures. [europepmc]
- Effectiveness and success factors of educational inhaler technique interventions in asthma & COPD patients: a systematic review. [europepmc]
- A Deep Learning-Based Radiomics Model for Prediction of Survival in Glioblastoma Multiforme. [europepmc]
- A Systematic Review of Studies Comparing the Measurement Properties of the Three-Level and Five-Level Versions of the EQ-5D. [europepmc]
- Análise de concordância em estudos clínicos e experimentais. [europepmc]
- The usefulness of cardiac CT in the diagnosis of perivalvular complications in patients with infective endocarditis. [europepmc]
- Deep Learning-Assisted Diagnosis of Cerebral Aneurysms Using the HeadXNet Model. [europepmc]
- The superior predictive value of 166 Ho-scout compared with 99m Tc-macroaggregated albumin prior to 166 Ho-microspheres radioembolization in patients with liver metastases. [europepmc]
- Comparability of three intraocular pressure measurement: iCare pro rebound, non-contact and Goldmann applanation tonometry in different IOP group. [europepmc]
- CO-RADS: A Categorical CT Assessment Scheme for Patients Suspected of Having COVID-19-Definition and Evaluation. [europepmc]
- A Step-by-Step Process on Sample Size Determination for Medical Research. [europepmc]
- Kappa statistic considerations in evaluating inter-rater reliability between two raters: which, when and context matters. [europepmc]